robots.txt & llms.txt across the Tranco top 1M
Data as of 2026-07-25 · robots.txt via Common Crawl CC-MAIN-2026-25 + our polite crawl · llms.txt via our polite crawl · Tranco top-1,000,000 · crawls: 2026-07-25 · 2026-07-17 687,794 robots.txt fetches · 491,025 parsed · 54,000 llms.txt probes

AI crawlers

Share of parsed robots.txt files across the top 1M that name each AI user-agent in a directive — a targeting rate (the bot is addressed by a rule), not an allow-vs-deny split (the data records that a bot is addressed, not the decision). Bars are tinted by vendor.

Weighting: Raw Rank-weighted
directive tokens (Google-Extended, Applebot-Extended) are robots.txt opt-out directives, NOT crawlers that appear in request logs. We model declared opt-out separately from actual-traffic presence (PRD FR-17b).

Mixed signals: llms.txt welcome mat vs robots.txt block

The flagship cross-tab: does a site's llms.txt (a welcome mat for LLMs) agree with its robots.txt AI-crawler rules? The mixed cell — publishes an llms.txt and blocks AI crawlers — is a genuine contradiction.

data table
quadrantdomains
welcoming (llms.txt + allows AI)631
mixed (llms.txt + blocks AI)560
blocking (blocks AI, no llms.txt)56276
indifferent (neither)14862

Domains with a parsed robots.txt this crawl, crossed with genuine (200, non-HTML) llms.txt presence. The mixed cell is the headline contradiction.

Targeting leaderboard

Anthropic 63.4% OpenAI 57.8% Google 32.9% Perplexity 30.2% Apple 24.3% Meta 23.8%
ai2bot · 82.5% blocked 4.5%
ai2bot-dolma · 88.2% blocked 3.3%
semanticscholarbot · 84.6% blocked 0.1%
amazonbot · 89.1% blocked 15.1%
claudebot · 79.0% blocked 21.6%
anthropic-ai · 63.8% blocked 15.2%
claude-web · 64.1% blocked 14.2%
claude-searchbot · 42.2% blocked 6.3%
claude-user · 36.9% blocked 6.3%
applebot-extended directive · 88.9% blocked 15.5%
applebot · 31.8% blocked 8.8%
bytespider · 89.7% blocked 13.8%
tiktokspider · 82.5% blocked 1.2%
cohere-ai · 69.5% blocked 13.8%
cohere-training-data-crawler · 86.3% blocked 3.7%
ccbot · 85.8% blocked 22.8%
deepseekbot · 74.5% blocked 2.8%
diffbot · 82.1% blocked 6.6%
duckassistbot · 55.9% blocked 4.8%
google-extended directive · 80.2% blocked 29.6%
googleother · 47.0% blocked 3.3%
petalbot · 80.6% blocked 11.6%
img2dataset · 89.2% blocked 4.0%
kangaroo bot · 88.1% blocked 3.1%
meta-externalagent · 87.1% blocked 12.0%
facebookbot · 68.2% blocked 7.4%
meta-externalfetcher · 65.0% blocked 4.3%
mistralai-user · 66.1% blocked 3.7%
gptbot · 78.3% blocked 30.8%
chatgpt-user · 47.8% blocked 17.7%
oai-searchbot · 32.4% blocked 9.2%
panscient.com · 95.5% blocked 2.0%
perplexitybot · 39.3% blocked 24.7%
perplexity-user · 34.1% blocked 5.5%
scrapy · 88.2% blocked 9.7%
timpibot · 88.6% blocked 4.6%
imgproxy · 98.0% blocked 1.4%
omgili · 91.0% blocked 6.4%
omgilibot · 89.1% blocked 4.7%
webzio-extended · 91.6% blocked 3.9%
youbot · 67.0% blocked 5.6%

Prior-art anchor (HTTP Almanac 2025, broad-web home-page census): GPTBot named on ~4.5% of files, up from ~2.9% the prior year — a near-doubling. Our top-1M numbers skew higher because popular sites have more (and more complex) robots.txt.