robots.txt & llms.txt across the Tranco top 1M
Data as of 2026-07-25 · robots.txt via Common Crawl CC-MAIN-2026-25 + our polite crawl · llms.txt via our polite crawl · Tranco top-1,000,000 · crawls: 2026-07-25 · 2026-07-17 687,794 robots.txt fetches · 491,025 parsed · 54,000 llms.txt probes

The State of robots.txt & llms.txt

A written summary of what Bot For You measures across the Tranco top 1M, placed alongside the authoritative broad-web baselines from the HTTP Archive Web Almanac 2025 (PRD §12.3). Every number here is measured, not modelled.

Preliminary — sample, not web-wide. Preliminary — sample, not web-wide. These figures come from a small, top-rank-skewed sample of the Tranco top 1M (popular sites over-index on having robots.txt/llms.txt), so adoption and prevalence read HIGHER than a broad-web census. Do not cite as definitive web-wide numbers. Treat the figures below as a directional read of the current sample, and cite the Almanac baselines (not our numbers) as web-wide facts until the full-1M crawl lands.

This crawl: 687,794 robots.txt fetches, 491,025 parsed, 54,000 llms.txt probes.

Headline findings vs prior art

MetricOur sampleAlmanac 2025 (web-wide)Status
robots.txt returns HTTP 200 sample 72.1% 84.9% Popular sites almost always serve a robots.txt, so our top-rank sample runs above the ~85% broad-web presence baseline.
robots.txt absent (404) sample 8.3% 13.0% Mirror of presence — the sample under-counts absence vs the ~13% broad-web 404 rate.
wildcard `*` user-agent present sample 94.3% 77.0% Almanac 2025 baseline ~77%.
llms.txt adoption (published) sample 11.1% 2.1% STRONGLY sample-skewed: Almanac 2025 measured ~2.1% web-wide; our top-rank sample reads far higher. This is adoption (published), NOT consumption — prior art (Ahrefs) found ~97% of published files are never read by an AI bot, and adoption is largely CMS/plugin-driven.
disallow-all (blocks every crawler) sample 2.7% No clean web-wide baseline; interpret directionally within the sample only.
allow-all (effectively fully open) sample 25.7% Sample figure only.
declares a Sitemap sample 63.7% Sample figure only.

Every row is tagged sample because our current corpus is small and top-rank-skewed; popular sites over-index on having (and elaborating) robots.txt/llms.txt. Do not publish these as definitive web-wide numbers.

robots.txt: presence and posture

In our sample, 72.1% of Tranco domains serve a robots.txt with HTTP 200 — higher than the Almanac's broad-web 84.9%, exactly as expected for a rank-skewed sample (popular sites almost always publish one). The mirror figure, absence (404), reads 8.3% here vs the Almanac's 13.0%. Because the RFC 9309 status semantics matter for interpretation — a 4xx means allow-all, a persistent 5xx means treat-as-disallow — the effective Google posture is reported separately on the dashboard rather than inferred from status codes alone.

llms.txt: adoption is real, consumption is not

We measure adoption (a site publishes an llms.txt) and conformance (it matches the llmstxt.org structure) — never consumption (an AI actually reads it).

Our sample reads llms.txt adoption at 11.1%. This is far above the Almanac's web-wide 2.1% — and the gap is almost entirely the sample skew, not a genuine web-wide surge. Prior art sharpens the honesty here: Ahrefs (137K sites) found ~97% of published llms.txt files are never fetched by an AI bot, AI crawlers don't probe for missing ones, and the Almanac attributes ~39.6% of valid files to the All in One SEO WordPress plugin — i.e. adoption is largely a passive CMS-plugin default, not deliberate strategy. Of the files we did find, 88.7% conform to the structure spec. See the llms.txt page for the full skeptic evidence.

AI-crawler targeting

Share of parsed robots.txt files that name each AI user-agent in a directive (a targeting rate — the bot is addressed — not an allow/deny split). directive marks opt-out tokens (e.g. Google-Extended) that aren't real request-log crawlers. Top 5 shown — see the full vendor-grouped leaderboard on the AI crawlers page.

AI user-agenttargeting rate (sample)
gptbot 21.1%
oai-searchbot 4.9%
chatgpt-user 7.2%
claudebot 18.5%
claude-user 1.7%

→ Full AI-crawler leaderboard (all vendors)

For scale, the Almanac's web-wide GPTBot mention rate was ~4.5% in 2025 (near-doubling year over year). Our sample rate runs higher for the same rank-skew reason. Inference-derived stats elsewhere on the site are labelled hypothesis; the targeting rates above are measured directly from parsed directives.

How to cite

Until the full Tranco-1M crawl lands, cite the Almanac 2025 figures as the web-wide baseline and describe Bot For You's numbers as "a preliminary top-rank sample." Methodology, the pinned Tranco list + Common Crawl crawl IDs, and downloadable exports are on the Methodology & Open Data page.

Bot For You (botfy.com). "The State of robots.txt & llms.txt." Tranco top-1,000,000, Common Crawl CC-MAIN-2026-25, snapshot 2026-07-25. Retrieved from https://botfy.com/report