robots.txt & llms.txt across the Tranco top 1M
Data as of 2026-09-07 · robots.txt via Common Crawl CC-MAIN-2026-34 + our polite crawl · llms.txt via our polite crawl · Tranco list XN67N (2026-09-03T22:00:02.533886) · 1,000,000 panel domains · crawls: 2026-09-07 · 2026-08-02 · 2026-07-25 · 2026-07-17 1,212,578 robots.txt fetches · 740,113 parsed · 56,498 llms.txt probes

Which AI crawlers do websites block?

Share of parsed robots.txt files across the top 1M that name each AI user-agent in a directive — a targeting rate (the bot is addressed by a rule), not an allow-vs-deny split (the data records that a bot is addressed, not the decision). Bars are tinted by vendor.

Weighting: Raw Rank-weighted
directive tokens (Google-Extended, Applebot-Extended) are robots.txt opt-out directives, NOT crawlers that appear in request logs. We model declared opt-out separately from actual-traffic presence (PRD FR-17b).

Related: which sites let AI crawlers in while blocking Google — the inverse configuration, and rarer than it looks.

Do sites that publish llms.txt also block AI crawlers?

The flagship cross-tab: does a site's llms.txt (a welcome mat for LLMs) agree with its robots.txt AI-crawler rules? The mixed cell — publishes an llms.txt and blocks AI crawlers — is a genuine contradiction.

data table
quadrantdomains
welcoming (llms.txt + allows AI)932
mixed (llms.txt + blocks AI)745
blocking (blocks AI, no llms.txt)128169
indifferent (neither)73463

Domains with a parsed robots.txt this crawl, crossed with genuine llms.txt presence (a 200 whose body is a real document). The mixed cell is the headline contradiction.

Which AI crawlers are named most often?

Anthropic 51.3% OpenAI 46.0% Google 26.5% Perplexity 23.9% Apple 22.9% Meta 21.5%
ai2bot · 63.9% blocked 2.9%
ai2bot-dolma · 59.8% blocked 2.1%
semanticscholarbot · 90.2% blocked 0.1%
amazonbot · 90.1% blocked 13.7%
claudebot · 83.4% blocked 18.4%
anthropic-ai · 62.5% blocked 11.7%
claude-web · 63.9% blocked 11.0%
claude-searchbot · 40.3% blocked 5.1%
claude-user · 34.4% blocked 5.1%
applebot-extended directive · 90.3% blocked 13.9%
applebot · 7.3% blocked 9.0%
bytespider · 63.9% blocked 12.9%
tiktokspider · 79.3% blocked 0.8%
cohere-ai · 67.2% blocked 10.8%
cohere-training-data-crawler · 82.9% blocked 2.4%
ccbot · 63.3% blocked 19.5%
deepseekbot · 79.7% blocked 1.9%
diffbot · 72.2% blocked 5.2%
duckassistbot · 55.6% blocked 4.0%
google-extended directive · 84.9% blocked 24.2%
googleother · 34.6% blocked 2.3%
petalbot · 26.9% blocked 10.1%
img2dataset · 62.1% blocked 2.5%
kangaroo bot · 55.2% blocked 2.0%
meta-externalagent · 87.6% blocked 11.8%
facebookbot · 61.7% blocked 5.0%
meta-externalfetcher · 49.6% blocked 4.7%
mistralai-user · 62.3% blocked 2.5%
gptbot · 82.8% blocked 25.1%
chatgpt-user · 49.5% blocked 13.7%
oai-searchbot · 26.9% blocked 7.2%
panscient.com · 95.8% blocked 1.2%
perplexitybot · 42.2% blocked 19.2%
perplexity-user · 33.1% blocked 4.6%
scrapy · 73.3% blocked 8.5%
timpibot · 72.5% blocked 3.7%
imgproxy · 98.8% blocked 0.9%
omgili · 81.4% blocked 5.2%
omgilibot · 80.4% blocked 4.2%
webzio-extended · 65.6% blocked 2.6%
youbot · 65.9% blocked 4.0%

Prior-art anchor (HTTP Almanac 2025, broad-web home-page census): GPTBot named on ~4.5% of files, up from ~2.9% the prior year — a near-doubling. Our top-1M numbers skew higher because popular sites have more (and more complex) robots.txt.