robots.txt & llms.txt across the Tranco top 1M
Data as of 2026-07-25 · robots.txt via Common Crawl CC-MAIN-2026-25 + our polite crawl · llms.txt via our polite crawl · Tranco top-1,000,000 · crawls: 2026-07-25 · 2026-07-17 687,794 robots.txt fetches · 491,025 parsed · 54,000 llms.txt probes

AI crawlers

Share of parsed robots.txt files across the top 1M that name each AI user-agent in a directive — a targeting rate (the bot is addressed by a rule), not an allow-vs-deny split (the data records that a bot is addressed, not the decision). Bars are tinted by vendor.

Weighting: Raw Rank-weighted
directive tokens (Google-Extended, Applebot-Extended) are robots.txt opt-out directives, NOT crawlers that appear in request logs. We model declared opt-out separately from actual-traffic presence (PRD FR-17b).

Mixed signals: llms.txt welcome mat vs robots.txt block

The flagship cross-tab: does a site's llms.txt (a welcome mat for LLMs) agree with its robots.txt AI-crawler rules? The mixed cell — publishes an llms.txt and blocks AI crawlers — is a genuine contradiction.

data table
quadrantdomains
welcoming (llms.txt + allows AI)631
mixed (llms.txt + blocks AI)560
blocking (blocks AI, no llms.txt)56276
indifferent (neither)14862

Domains with a parsed robots.txt this crawl, crossed with genuine (200, non-HTML) llms.txt presence. The mixed cell is the headline contradiction.

Targeting leaderboard

OpenAI 33.2% Anthropic 31.6% Meta 21.2% Common Crawl 18.3% Google 18.0% Apple 17.7%
ai2bot · 82.5% blocked 1.9%
ai2bot-dolma · 88.2% blocked 1.4%
semanticscholarbot · 84.6% blocked 0.1%
amazonbot · 89.1% blocked 16.3%
claudebot · 79.0% blocked 18.5%
anthropic-ai · 63.8% blocked 5.4%
claude-web · 64.1% blocked 4.1%
claude-searchbot · 42.2% blocked 1.9%
claude-user · 36.9% blocked 1.7%
applebot-extended directive · 88.9% blocked 14.1%
applebot · 31.8% blocked 3.6%
bytespider · 89.7% blocked 16.6%
tiktokspider · 82.5% blocked 0.8%
cohere-ai · 69.5% blocked 3.5%
cohere-training-data-crawler · 86.3% blocked 1.0%
ccbot · 85.8% blocked 18.3%
deepseekbot · 74.5% blocked 0.5%
diffbot · 82.1% blocked 3.0%
duckassistbot · 55.9% blocked 1.3%
google-extended directive · 80.2% blocked 16.6%
googleother · 47.0% blocked 1.4%
petalbot · 80.6% blocked 6.1%
img2dataset · 89.2% blocked 1.5%
kangaroo bot · 88.1% blocked 1.0%
meta-externalagent · 87.1% blocked 15.8%
facebookbot · 68.2% blocked 3.7%
meta-externalfetcher · 65.0% blocked 1.7%
mistralai-user · 66.1% blocked 0.9%
gptbot · 78.3% blocked 21.1%
chatgpt-user · 47.8% blocked 7.2%
oai-searchbot · 32.4% blocked 4.9%
panscient.com · 95.5% blocked 0.4%
perplexitybot · 39.3% blocked 6.7%
perplexity-user · 34.1% blocked 1.8%
scrapy · 88.2% blocked 2.3%
timpibot · 88.6% blocked 2.1%
imgproxy · 98.0% blocked 0.1%
omgili · 91.0% blocked 2.7%
omgilibot · 89.1% blocked 2.7%
webzio-extended · 91.6% blocked 1.5%
youbot · 67.0% blocked 2.9%

Prior-art anchor (HTTP Almanac 2025, broad-web home-page census): GPTBot named on ~4.5% of files, up from ~2.9% the prior year — a near-doubling. Our top-1M numbers skew higher because popular sites have more (and more complex) robots.txt.