robots.txt & llms.txt across the Tranco top 1M
Data as of 2026-09-07 · robots.txt via Common Crawl CC-MAIN-2026-34 + our polite crawl · llms.txt via our polite crawl · Tranco list XN67N (2026-09-03T22:00:02.533886) · 1,000,000 panel domains · crawls: 2026-09-07 · 2026-08-02 · 2026-07-25 · 2026-07-17 1,212,578 robots.txt fetches · 740,113 parsed · 56,498 llms.txt probes

Which AI crawlers do websites block?

Share of parsed robots.txt files across the top 1M that name each AI user-agent in a directive — a targeting rate (the bot is addressed by a rule), not an allow-vs-deny split (the data records that a bot is addressed, not the decision). Bars are tinted by vendor.

Weighting: Raw Rank-weighted
directive tokens (Google-Extended, Applebot-Extended) are robots.txt opt-out directives, NOT crawlers that appear in request logs. We model declared opt-out separately from actual-traffic presence (PRD FR-17b).

Related: which sites let AI crawlers in while blocking Google — the inverse configuration, and rarer than it looks.

Do sites that publish llms.txt also block AI crawlers?

The flagship cross-tab: does a site's llms.txt (a welcome mat for LLMs) agree with its robots.txt AI-crawler rules? The mixed cell — publishes an llms.txt and blocks AI crawlers — is a genuine contradiction.

data table
quadrantdomains
welcoming (llms.txt + allows AI)932
mixed (llms.txt + blocks AI)745
blocking (blocks AI, no llms.txt)128169
indifferent (neither)73463

Domains with a parsed robots.txt this crawl, crossed with genuine llms.txt presence (a 200 whose body is a real document). The mixed cell is the headline contradiction.

Which AI crawlers are named most often?

Anthropic 25.1% OpenAI 25.0% Common Crawl 21.9% Apple 21.9% ByteDance 21.4% Meta 17.4%
ai2bot · 63.9% blocked 1.3%
ai2bot-dolma · 59.8% blocked 0.9%
semanticscholarbot · 90.2% blocked 0.1%
amazonbot · 90.1% blocked 14.6%
claudebot · 83.4% blocked 15.8%
anthropic-ai · 62.5% blocked 3.8%
claude-web · 63.9% blocked 2.9%
claude-searchbot · 40.3% blocked 1.4%
claude-user · 34.4% blocked 1.2%
applebot-extended directive · 90.3% blocked 13.3%
applebot · 7.3% blocked 8.6%
bytespider · 63.9% blocked 20.9%
tiktokspider · 79.3% blocked 0.5%
cohere-ai · 67.2% blocked 2.7%
cohere-training-data-crawler · 82.9% blocked 0.6%
ccbot · 63.3% blocked 21.9%
deepseekbot · 79.7% blocked 0.3%
diffbot · 72.2% blocked 1.9%
duckassistbot · 55.6% blocked 1.0%
google-extended directive · 84.9% blocked 14.6%
googleother · 34.6% blocked 1.0%
petalbot · 26.9% blocked 9.8%
img2dataset · 62.1% blocked 0.9%
kangaroo bot · 55.2% blocked 0.7%
meta-externalagent · 87.6% blocked 13.9%
facebookbot · 61.7% blocked 2.3%
meta-externalfetcher · 49.6% blocked 1.2%
mistralai-user · 62.3% blocked 0.6%
gptbot · 82.8% blocked 17.0%
chatgpt-user · 49.5% blocked 4.8%
oai-searchbot · 26.9% blocked 3.2%
panscient.com · 95.8% blocked 0.2%
perplexitybot · 42.2% blocked 4.6%
perplexity-user · 33.1% blocked 1.2%
scrapy · 73.3% blocked 1.4%
timpibot · 72.5% blocked 1.3%
imgproxy · 98.8% blocked 0.1%
omgilibot · 80.4% blocked 2.2%
omgili · 81.4% blocked 2.2%
webzio-extended · 65.6% blocked 1.0%
youbot · 65.9% blocked 2.3%

Prior-art anchor (HTTP Almanac 2025, broad-web home-page census): GPTBot named on ~4.5% of files, up from ~2.9% the prior year — a near-doubling. Our top-1M numbers skew higher because popular sites have more (and more complex) robots.txt.