robots.txt & llms.txt across the Tranco top 1M
Data as of 2026-07-17 · robots.txt via Common Crawl CC-MAIN-2026-25 + our polite crawl · llms.txt via our polite crawl · Tranco top-1,000,000 · crawls: 2026-07-25 · 2026-07-17 746,000 robots.txt fetches · 574,664 parsed · 54,000 llms.txt probes

AI crawlers

Share of parsed robots.txt files across the top 1M that name each AI user-agent in a directive — a targeting rate (the bot is addressed by a rule), not an allow-vs-deny split (the data records that a bot is addressed, not the decision). Bars are tinted by vendor.

Weighting: Raw Rank-weighted
directive tokens (Google-Extended, Applebot-Extended) are robots.txt opt-out directives, NOT crawlers that appear in request logs. We model declared opt-out separately from actual-traffic presence (PRD FR-17b).

Mixed signals: llms.txt welcome mat vs robots.txt block

The flagship cross-tab: does a site's llms.txt (a welcome mat for LLMs) agree with its robots.txt AI-crawler rules? The mixed cell — publishes an llms.txt and blocks AI crawlers — is a genuine contradiction.

data table
quadrantdomains
welcoming (llms.txt + allows AI)478
mixed (llms.txt + blocks AI)484
blocking (blocks AI, no llms.txt)52988
indifferent (neither)13224

Domains with a parsed robots.txt this crawl, crossed with genuine (200, non-HTML) llms.txt presence. The mixed cell is the headline contradiction.

Targeting leaderboard

Anthropic 50.6% OpenAI 47.9% Google 25.3% Apple 24.5% Perplexity 24.4% Common Crawl 18.6%
ai2bot · 84.3% blocked 2.8%
ai2bot-dolma · 90.3% blocked 1.9%
semanticscholarbot · 84.3% blocked 0.1%
amazonbot · 90.0% blocked 15.3%
claudebot · 80.7% blocked 19.1%
anthropic-ai · 65.3% blocked 12.2%
claude-web · 65.8% blocked 10.9%
claude-searchbot · 43.3% blocked 4.4%
claude-user · 37.6% blocked 4.0%
applebot-extended directive · 90.2% blocked 13.9%
applebot · 31.8% blocked 10.6%
bytespider · 90.1% blocked 11.0%
tiktokspider · 82.7% blocked 0.9%
cohere-ai · 71.4% blocked 11.6%
cohere-training-data-crawler · 86.8% blocked 2.1%
ccbot · 86.7% blocked 18.6%
deepseekbot · 74.3% blocked 1.3%
diffbot · 83.0% blocked 4.8%
duckassistbot · 56.9% blocked 3.5%
google-extended directive · 82.0% blocked 23.3%
googleother · 47.9% blocked 2.1%
petalbot · 80.4% blocked 11.7%
img2dataset · 91.2% blocked 2.6%
kangaroo bot · 91.3% blocked 1.8%
meta-externalagent · 87.9% blocked 9.7%
facebookbot · 69.2% blocked 5.3%
meta-externalfetcher · 66.8% blocked 3.5%
mistralai-user · 67.6% blocked 1.9%
gptbot · 79.9% blocked 26.4%
chatgpt-user · 49.3% blocked 14.1%
oai-searchbot · 33.1% blocked 7.4%
panscient.com · 95.6% blocked 1.0%
perplexitybot · 40.6% blocked 21.1%
perplexity-user · 34.5% blocked 3.3%
scrapy · 89.6% blocked 9.6%
timpibot · 89.9% blocked 3.3%
imgproxy · 98.1% blocked 0.4%
omgili · 92.1% blocked 4.9%
omgilibot · 90.2% blocked 4.3%
webzio-extended · 93.4% blocked 2.3%
youbot · 68.9% blocked 3.6%

Prior-art anchor (HTTP Almanac 2025, broad-web home-page census): GPTBot named on ~4.5% of files, up from ~2.9% the prior year — a near-doubling. Our top-1M numbers skew higher because popular sites have more (and more complex) robots.txt.