The State of robots.txt & llms.txt
A written summary of what Bot For You measures across the Tranco top 1M, placed alongside the authoritative broad-web baselines from the HTTP Archive Web Almanac 2025 (PRD §12.3). Every number here is measured, not modelled.
This crawl: 687,794 robots.txt fetches, 491,025 parsed, 54,000 llms.txt probes.
Headline findings vs prior art ¶
| Metric | Our sample | Almanac 2025 (web-wide) | Status |
|---|---|---|---|
| robots.txt returns HTTP 200 sample | 72.1% | 84.9% | Popular sites almost always serve a robots.txt, so our top-rank sample runs above the ~85% broad-web presence baseline. |
| robots.txt absent (404) sample | 8.3% | 13.0% | Mirror of presence — the sample under-counts absence vs the ~13% broad-web 404 rate. |
| wildcard `*` user-agent present sample | 94.3% | 77.0% | Almanac 2025 baseline ~77%. |
| llms.txt adoption (published) sample | 11.1% | 2.1% | STRONGLY sample-skewed: Almanac 2025 measured ~2.1% web-wide; our top-rank sample reads far higher. This is adoption (published), NOT consumption — prior art (Ahrefs) found ~97% of published files are never read by an AI bot, and adoption is largely CMS/plugin-driven. |
| disallow-all (blocks every crawler) sample | 2.7% | — | No clean web-wide baseline; interpret directionally within the sample only. |
| allow-all (effectively fully open) sample | 25.7% | — | Sample figure only. |
| declares a Sitemap sample | 63.7% | — | Sample figure only. |
Every row is tagged sample because our current corpus is small and top-rank-skewed; popular sites over-index on having (and elaborating) robots.txt/llms.txt. Do not publish these as definitive web-wide numbers.
robots.txt: presence and posture ¶
In our sample, 72.1% of Tranco domains serve a robots.txt with HTTP 200 — higher than the Almanac's broad-web 84.9%, exactly as expected for a rank-skewed sample (popular sites almost always publish one). The mirror figure, absence (404), reads 8.3% here vs the Almanac's 13.0%. Because the RFC 9309 status semantics matter for interpretation — a 4xx means allow-all, a persistent 5xx means treat-as-disallow — the effective Google posture is reported separately on the dashboard rather than inferred from status codes alone.
llms.txt: adoption is real, consumption is not ¶
Our sample reads llms.txt adoption at 11.1%. This is far above the Almanac's web-wide 2.1% — and the gap is almost entirely the sample skew, not a genuine web-wide surge. Prior art sharpens the honesty here: Ahrefs (137K sites) found ~97% of published llms.txt files are never fetched by an AI bot, AI crawlers don't probe for missing ones, and the Almanac attributes ~39.6% of valid files to the All in One SEO WordPress plugin — i.e. adoption is largely a passive CMS-plugin default, not deliberate strategy. Of the files we did find, 88.7% conform to the structure spec. See the llms.txt page for the full skeptic evidence.
AI-crawler targeting ¶
Share of parsed robots.txt files that name each AI user-agent in a directive (a targeting rate — the bot is addressed — not an allow/deny split). directive marks opt-out tokens (e.g. Google-Extended) that aren't real request-log crawlers. Top 5 shown — see the full vendor-grouped leaderboard on the AI crawlers page.
| AI user-agent | targeting rate (sample) |
|---|---|
| gptbot | 21.1% |
| oai-searchbot | 4.9% |
| chatgpt-user | 7.2% |
| claudebot | 18.5% |
| claude-user | 1.7% |
→ Full AI-crawler leaderboard (all vendors)
For scale, the Almanac's web-wide GPTBot mention rate was ~4.5% in 2025 (near-doubling year over year). Our sample rate runs higher for the same rank-skew reason. Inference-derived stats elsewhere on the site are labelled hypothesis; the targeting rates above are measured directly from parsed directives.
How to cite ¶
Until the full Tranco-1M crawl lands, cite the Almanac 2025 figures as the web-wide baseline and describe Bot For You's numbers as "a preliminary top-rank sample." Methodology, the pinned Tranco list + Common Crawl crawl IDs, and downloadable exports are on the Methodology & Open Data page.
Bot For You (botfy.com). "The State of robots.txt & llms.txt." Tranco top-1,000,000, Common Crawl CC-MAIN-2026-25, snapshot 2026-07-25. Retrieved from https://botfy.com/report