The State of robots.txt & llms.txt
A written summary of what Bot For You measures across the Tranco top 1M, placed alongside the authoritative broad-web baselines from the HTTP Archive Web Almanac 2025 (PRD §12.3). Every number here is measured, not modelled.
This crawl: 1,212,578 robots.txt fetches, 740,113 parsed; llms.txt balanced panel: 50,000 selected, 30,676 live.
Headline findings vs prior art ¶
| Metric | Our result | Population / denominator | Almanac 2025 (broad web) | Interpretation |
|---|---|---|---|---|
| robots.txt returns HTTP 200 ranked panel | 61.7% | Tranco-ranked robots.txt fetches n=1,212,578 |
84.9% | Bot For You measures the ranked web; Almanac measures a broader web population, so the rates are context rather than like-for-like. |
| robots.txt absent (404) ranked panel | 8.4% | Tranco-ranked robots.txt fetches n=1,212,578 |
13.0% | A real HTTP 404 only; fetch errors and throttles are not counted as absence. |
| wildcard `*` user-agent present parsed files | 89.8% | successfully parsed robots.txt files n=740,113 |
77.0% | Conditional on a fetched file that Bot For You could parse; Almanac's broad-web population is shown only as context. |
| llms.txt adoption (published) balanced cohort | 13.0% | strict-200 live sites in the balanced top-500k panel n=30,676 |
2.1% | Balanced top-500k panel, among live sites; live-count-weighted across ten equal 50k rank bands. This is adoption (published), NOT consumption. |
| disallow-all (blocks every crawler) parsed files | 10.3% | successfully parsed robots.txt files n=740,113 |
— | No clean broad-web baseline; the denominator is parsed files, not all ranked domains. |
| allow-all (effectively fully open) parsed files | 25.7% | successfully parsed robots.txt files n=740,113 |
— | Conditional on parsed robots.txt files; not a rate over domains whose file was absent or unavailable. |
| declares a Sitemap parsed files | 60.7% | successfully parsed robots.txt files n=740,113 |
— | Conditional on parsed robots.txt files; no clean broad-web baseline is used. |
A percentage is only comparable when its population and denominator match. Bot For You's ranked-web figures should not be restated as estimates for the entire web, and the Almanac baseline should not be treated as if it measured this same Tranco panel.
robots.txt: presence and posture ¶
Across this snapshot's 1,212,578 Tranco-ranked robots.txt fetches, 61.7% returned HTTP 200 and 8.4% returned a real HTTP 404. The Almanac's broad-web figures are 84.9% and 13.0%, respectively. The populations differ: Bot For You measures the ranked web, while the Almanac measures a broader web corpus. This is a population distinction, not evidence that Bot For You is still waiting for a full Tranco crawl. Because RFC 9309 status semantics matter — a 4xx means allow-all, while a persistent 5xx means treat-as-disallow — effective Google posture is reported separately on the dashboard rather than inferred from status codes alone.
llms.txt: adoption is real, consumption is not ¶
Our balanced top-500k panel reads llms.txt adoption among live sites at 13.0%. This is live-count-weighted across ten equal 50k rank bands; the separate top-50k census is down-sampled to 5,000 before it contributes, so popular sites do not receive ten times the influence. The Almanac's broad-web 2.1% is useful context but not the same population or liveness-gated denominator. Prior art sharpens the honesty here: Ahrefs (137K sites) found ~97% of published llms.txt files are never fetched by an AI bot, AI crawlers don't probe for missing ones, and the Almanac attributes ~39.6% of valid files to the All in One SEO WordPress plugin — i.e. adoption is often a passive CMS-plugin default, not deliberate strategy. Of the files in the balanced panel, 96.9% conform to the structure spec. See the llms.txt page for the full skeptic evidence.
AI-crawler targeting ¶
Share of parsed robots.txt files that name each AI user-agent in a directive (a targeting rate — the bot is addressed — not an allow/deny split). directive marks opt-out tokens (e.g. Google-Extended) that aren't real request-log crawlers. Top 5 shown — see the full vendor-grouped leaderboard on the AI crawlers page.
| AI user-agent | targeting rate (parsed robots.txt, n=740,113) |
|---|---|
| gptbot | 17.0% |
| oai-searchbot | 3.2% |
| chatgpt-user | 4.8% |
| claudebot | 15.8% |
| claude-user | 1.2% |
→ Full AI-crawler leaderboard (all vendors)
For scale, the Almanac's web-wide GPTBot mention rate was ~4.5% in 2025 (near-doubling year over year). Bot For You's denominator is the 740,113 successfully parsed robots.txt files in its ranked panel, not the Almanac's broad-web population, so the two are context rather than a direct trend. Inference-derived stats elsewhere on the site are labelled hypothesis; the targeting rates above are measured directly from parsed directives.
How to cite ¶
Cite each figure with its population: robots.txt HTTP status over 1,212,578 Tranco-ranked fetches; directive and AI-targeting metrics over 740,113 parsed robots.txt files; and llms.txt adoption over 30,676 strict-200 live sites from a 50,000-domain, balanced top-500k panel. Do not describe these collectively as web-wide, but do not understate the completed robots crawl as incomplete either. Methodology, pinned source IDs, and downloads are on the Methodology & Open Data page.
Bot For You (botfy.com). "The State of robots.txt & llms.txt." , Common Crawl CC-MAIN-2026-34, snapshot 2026-09-07; robots.txt panel n=1212578; llms.txt balanced panel n=50000, live n=30676. Retrieved from https://botfy.com/report