robots.txt & llms.txt across the Tranco top 1M
Data as of 2026-07-17 · robots.txt via Common Crawl crawl ID not recorded for this snapshot + our polite crawl · llms.txt via our polite crawl · Tranco list ID not recorded for this snapshot · crawls: 2026-09-07 · 2026-08-02 · 2026-07-25 · 2026-07-17 746,000 robots.txt fetches · 574,664 parsed · 54,000 llms.txt probes

Does robots.txt practice change with site popularity?

Headline metrics broken down by segment. Toggle the segmentation axis and the weighting. Rank-weighting (FR-23b) weights each domain by 1/Tranco-rank, so a restriction on a top site counts far more than one deep in the tail.

One crawl ingested (2026-07-17). Cross-crawl trend lines appear automatically when the next monthly crawl lands.

Weighting: Raw Rank-weighted
robots.txt returns 200
data table
Top 1,00068.6%
Top 10,00072.3%
Top 100,00069.3%
Top 1,000,00068.1%
absent (404)
data table
Top 1,00012.4%
Top 10,00010.7%
Top 100,00012.6%
Top 1,000,00011.9%
disallow-all
data table
Top 1,00019.0%
Top 10,0008.4%
Top 100,0005.9%
Top 1,000,0004.3%
allow-all
data table
Top 1,00013.5%
Top 10,00016.7%
Top 100,00022.0%
Top 1,000,00026.0%
wildcard * UA
data table
Top 1,00094.5%
Top 10,00095.2%
Top 100,00094.5%
Top 1,000,00093.8%
declares a sitemap
data table
Top 1,00054.5%
Top 10,00058.5%
Top 100,00058.6%
Top 1,000,00059.7%

Full data table

Rank tier sample status 200status 404disallow allallow allwildcard uasitemap
Top 1,000 839 68.6% 12.4% 19.0% 13.5% 94.5% 54.5%
Top 10,000 7,252 72.3% 10.7% 8.4% 16.7% 95.2% 58.5%
Top 100,000 63,980 69.3% 12.6% 5.9% 22.0% 94.5% 58.6%
Top 1,000,000 502,593 68.1% 11.9% 4.3% 26.0% 93.8% 59.7%

Representative llms.txt prevalence uses a different, balanced ten-band cohort and is published separately on the llms.txt page. The generic observed-pool rate remains internal because its top-50k census would overweight the most popular band.