robots.txt & llms.txt across the Tranco top 1M
Data as of 2026-07-17 · robots.txt via Common Crawl CC-MAIN-2026-25 + our polite crawl · llms.txt via our polite crawl · Tranco top-1,000,000 · crawls: 2026-07-25 · 2026-07-17 746,000 robots.txt fetches · 574,664 parsed · 54,000 llms.txt probes

AI crawlers

Share of parsed robots.txt files across the top 1M that name each AI user-agent in a directive — a targeting rate (the bot is addressed by a rule), not an allow-vs-deny split (the data records that a bot is addressed, not the decision). Bars are tinted by vendor.

Weighting: Raw Rank-weighted
directive tokens (Google-Extended, Applebot-Extended) are robots.txt opt-out directives, NOT crawlers that appear in request logs. We model declared opt-out separately from actual-traffic presence (PRD FR-17b).

Mixed signals: llms.txt welcome mat vs robots.txt block

The flagship cross-tab: does a site's llms.txt (a welcome mat for LLMs) agree with its robots.txt AI-crawler rules? The mixed cell — publishes an llms.txt and blocks AI crawlers — is a genuine contradiction.

data table
quadrantdomains
welcoming (llms.txt + allows AI)478
mixed (llms.txt + blocks AI)484
blocking (blocks AI, no llms.txt)52988
indifferent (neither)13224

Domains with a parsed robots.txt this crawl, crossed with genuine (200, non-HTML) llms.txt presence. The mixed cell is the headline contradiction.

Targeting leaderboard

OpenAI 29.2% Anthropic 27.6% Meta 19.0% Apple 16.3% Common Crawl 16.2% Google 15.9%
ai2bot · 84.3% blocked 1.6%
ai2bot-dolma · 90.3% blocked 1.2%
semanticscholarbot · 84.3% blocked 0.1%
amazonbot · 90.0% blocked 14.8%
claudebot · 80.7% blocked 16.5%
anthropic-ai · 65.3% blocked 4.6%
claude-web · 65.8% blocked 3.5%
claude-searchbot · 43.3% blocked 1.6%
claude-user · 37.6% blocked 1.4%
applebot-extended directive · 90.2% blocked 12.6%
applebot · 31.8% blocked 3.8%
bytespider · 90.1% blocked 14.9%
tiktokspider · 82.7% blocked 0.7%
cohere-ai · 71.4% blocked 3.0%
cohere-training-data-crawler · 86.8% blocked 0.8%
ccbot · 86.7% blocked 16.2%
deepseekbot · 74.3% blocked 0.4%
diffbot · 83.0% blocked 2.6%
duckassistbot · 56.9% blocked 1.1%
google-extended directive · 82.0% blocked 14.7%
googleother · 47.9% blocked 1.2%
petalbot · 80.4% blocked 5.5%
img2dataset · 91.2% blocked 1.3%
kangaroo bot · 91.3% blocked 0.9%
meta-externalagent · 87.9% blocked 14.2%
facebookbot · 69.2% blocked 3.2%
meta-externalfetcher · 66.8% blocked 1.5%
mistralai-user · 67.6% blocked 0.7%
gptbot · 79.9% blocked 18.8%
chatgpt-user · 49.3% blocked 6.2%
oai-searchbot · 33.1% blocked 4.2%
panscient.com · 95.6% blocked 0.3%
perplexitybot · 40.6% blocked 5.6%
perplexity-user · 34.5% blocked 1.4%
scrapy · 89.6% blocked 2.0%
timpibot · 89.9% blocked 1.8%
imgproxy · 98.1% blocked 0.1%
omgili · 92.1% blocked 2.4%
omgilibot · 90.2% blocked 2.3%
webzio-extended · 93.4% blocked 1.3%
youbot · 68.9% blocked 2.4%

Prior-art anchor (HTTP Almanac 2025, broad-web home-page census): GPTBot named on ~4.5% of files, up from ~2.9% the prior year — a near-doubling. Our top-1M numbers skew higher because popular sites have more (and more complex) robots.txt.