robots.txt & llms.txt across the Tranco top 1M
Data as of 2026-07-25 · robots.txt via Common Crawl CC-MAIN-2026-25 + our polite crawl · llms.txt via our polite crawl · Tranco list ID not recorded for this snapshot · crawls: 2026-09-07 · 2026-08-02 · 2026-07-25 · 2026-07-17 687,794 robots.txt fetches · 491,025 parsed · 54,000 llms.txt probes

Does robots.txt practice change with site popularity?

Headline metrics broken down by segment. Toggle the segmentation axis and the weighting. Rank-weighting (FR-23b) weights each domain by 1/Tranco-rank, so a restriction on a top site counts far more than one deep in the tail.

One crawl ingested (2026-07-25). Cross-crawl trend lines appear automatically when the next monthly crawl lands.

Weighting: Raw Rank-weighted
robots.txt returns 200
data table
all72.1%
absent (404)
data table
all8.3%
disallow-all
data table
all2.7%
allow-all
data table
all25.7%
wildcard * UA
data table
all94.3%
declares a sitemap
data table
all63.7%

Full data table

Segment sample status 200status 404disallow allallow allwildcard uasitemap
all 491,025 72.1% 8.3% 2.7% 25.7% 94.3% 63.7%

Representative llms.txt prevalence uses a different, balanced ten-band cohort and is published separately on the llms.txt page. The generic observed-pool rate remains internal because its top-50k census would overweight the most popular band.