robots.txt & llms.txt across the Tranco top 1M
Data as of 2026-07-17 · robots.txt via Common Crawl crawl ID not recorded for this snapshot + our polite crawl · llms.txt via our polite crawl · Tranco list ID not recorded for this snapshot · crawls: 2026-09-07 · 2026-08-02 · 2026-07-25 · 2026-07-17 746,000 robots.txt fetches · 574,664 parsed · 54,000 llms.txt probes

The State of robots.txt & llms.txt

A written summary of what Bot For You measures across the Tranco top 1M, placed alongside the authoritative broad-web baselines from the HTTP Archive Web Almanac 2025 (PRD §12.3). Every number here is measured, not modelled.

Different questions use different denominators. HTTP-status rates use all measured Tranco-ranked robots.txt fetches; directive and AI-targeting rates use only files successfully parsed; llms.txt prevalence uses the strict-200 live population in its balanced top-500k cohort. HTTP Almanac covers the broader web, so its figures are context rather than a like-for-like denominator.

This crawl: 746,000 robots.txt fetches, 574,664 parsed; llms.txt balanced panel: 50,000 selected, 28,783 live.

Headline findings vs prior art

MetricOur resultPopulation / denominator Almanac 2025 (broad web)Interpretation
robots.txt returns HTTP 200 ranked panel 68.3% Tranco-ranked robots.txt fetches
n=746,000
84.9% Bot For You measures the ranked web; Almanac measures a broader web population, so the rates are context rather than like-for-like.
robots.txt absent (404) ranked panel 12.0% Tranco-ranked robots.txt fetches
n=746,000
13.0% A real HTTP 404 only; fetch errors and throttles are not counted as absence.
wildcard `*` user-agent present parsed files 93.9% successfully parsed robots.txt files
n=574,664
77.0% Conditional on a fetched file that Bot For You could parse; Almanac's broad-web population is shown only as context.
llms.txt adoption (published) balanced cohort 10.9% strict-200 live sites in the balanced top-500k panel
n=28,783
2.1% Balanced top-500k panel, among live sites; live-count-weighted across ten equal 50k rank bands. This is adoption (published), NOT consumption.
disallow-all (blocks every crawler) parsed files 4.5% successfully parsed robots.txt files
n=574,664
No clean broad-web baseline; the denominator is parsed files, not all ranked domains.
allow-all (effectively fully open) parsed files 25.5% successfully parsed robots.txt files
n=574,664
Conditional on parsed robots.txt files; not a rate over domains whose file was absent or unavailable.
declares a Sitemap parsed files 59.5% successfully parsed robots.txt files
n=574,664
Conditional on parsed robots.txt files; no clean broad-web baseline is used.

A percentage is only comparable when its population and denominator match. Bot For You's ranked-web figures should not be restated as estimates for the entire web, and the Almanac baseline should not be treated as if it measured this same Tranco panel.

robots.txt: presence and posture

Across this snapshot's 746,000 Tranco-ranked robots.txt fetches, 68.3% returned HTTP 200 and 12.0% returned a real HTTP 404. The Almanac's broad-web figures are 84.9% and 13.0%, respectively. The populations differ: Bot For You measures the ranked web, while the Almanac measures a broader web corpus. This is a population distinction, not evidence that Bot For You is still waiting for a full Tranco crawl. Because RFC 9309 status semantics matter — a 4xx means allow-all, while a persistent 5xx means treat-as-disallow — effective Google posture is reported separately on the dashboard rather than inferred from status codes alone.

llms.txt: adoption is real, consumption is not

We measure adoption (a site publishes an llms.txt) and conformance (it matches the llmstxt.org structure) — never consumption (an AI actually reads it).

Our balanced top-500k panel reads llms.txt adoption among live sites at 10.9%. This is live-count-weighted across ten equal 50k rank bands; the separate top-50k census is down-sampled to 5,000 before it contributes, so popular sites do not receive ten times the influence. The Almanac's broad-web 2.1% is useful context but not the same population or liveness-gated denominator. Prior art sharpens the honesty here: Ahrefs (137K sites) found ~97% of published llms.txt files are never fetched by an AI bot, AI crawlers don't probe for missing ones, and the Almanac attributes ~39.6% of valid files to the All in One SEO WordPress plugin — i.e. adoption is often a passive CMS-plugin default, not deliberate strategy. Of the files in the balanced panel, 89.6% conform to the structure spec. See the llms.txt page for the full skeptic evidence.

AI-crawler targeting

Share of parsed robots.txt files that name each AI user-agent in a directive (a targeting rate — the bot is addressed — not an allow/deny split). directive marks opt-out tokens (e.g. Google-Extended) that aren't real request-log crawlers. Top 5 shown — see the full vendor-grouped leaderboard on the AI crawlers page.

AI user-agent targeting rate (parsed robots.txt, n=574,664)
gptbot 18.8%
oai-searchbot 4.2%
chatgpt-user 6.2%
claudebot 16.5%
claude-user 1.4%

→ Full AI-crawler leaderboard (all vendors)

For scale, the Almanac's web-wide GPTBot mention rate was ~4.5% in 2025 (near-doubling year over year). Bot For You's denominator is the 574,664 successfully parsed robots.txt files in its ranked panel, not the Almanac's broad-web population, so the two are context rather than a direct trend. Inference-derived stats elsewhere on the site are labelled hypothesis; the targeting rates above are measured directly from parsed directives.

How to cite

Cite each figure with its population: robots.txt HTTP status over 746,000 Tranco-ranked fetches; directive and AI-targeting metrics over 574,664 parsed robots.txt files; and llms.txt adoption over 28,783 strict-200 live sites from a 50,000-domain, balanced top-500k panel. Do not describe these collectively as web-wide, but do not understate the completed robots crawl as incomplete either. Methodology, pinned source IDs, and downloads are on the Methodology & Open Data page.

Bot For You (botfy.com). "The State of robots.txt & llms.txt." , Common Crawl None, snapshot 2026-07-17; robots.txt panel n=746000; llms.txt balanced panel n=50000, live n=28783. Retrieved from https://botfy.com/report