robots.txt & llms.txt across the Tranco top 1M
Data as of 2026-08-02 · robots.txt via Common Crawl CC-MAIN-2026-30 + our polite crawl · llms.txt via our polite crawl · Tranco list L5684 (2026-07-14) · 1,000,000 panel domains · crawls: 2026-09-07 · 2026-08-02 · 2026-07-25 · 2026-07-17 1,000,000 robots.txt fetches · 608,834 parsed · 55,770 llms.txt probes

Which AI crawlers do websites block?

Share of parsed robots.txt files across the top 1M that name each AI user-agent in a directive — a targeting rate (the bot is addressed by a rule), not an allow-vs-deny split (the data records that a bot is addressed, not the decision). Bars are tinted by vendor.

Weighting: Raw Rank-weighted
directive tokens (Google-Extended, Applebot-Extended) are robots.txt opt-out directives, NOT crawlers that appear in request logs. We model declared opt-out separately from actual-traffic presence (PRD FR-17b).

Related: which sites let AI crawlers in while blocking Google — the inverse configuration, and rarer than it looks.

Do sites that publish llms.txt also block AI crawlers?

The flagship cross-tab: does a site's llms.txt (a welcome mat for LLMs) agree with its robots.txt AI-crawler rules? The mixed cell — publishes an llms.txt and blocks AI crawlers — is a genuine contradiction.

data table
quadrantdomains
welcoming (llms.txt + allows AI)785
mixed (llms.txt + blocks AI)603
blocking (blocks AI, no llms.txt)97097
indifferent (neither)48376

Domains with a parsed robots.txt this crawl, crossed with genuine llms.txt presence (a 200 whose body is a real document). The mixed cell is the headline contradiction.

Which AI crawlers are named most often?

OpenAI 23.3% Anthropic 23.2% Common Crawl 18.6% Apple 18.4% ByteDance 17.9% Meta 15.8%
ai2bot · 67.8% blocked 1.3%
ai2bot-dolma · 66.7% blocked 0.9%
semanticscholarbot · 92.2% blocked 0.1%
amazonbot · 89.6% blocked 12.9%
claudebot · 82.4% blocked 14.2%
anthropic-ai · 64.2% blocked 3.8%
claude-web · 65.9% blocked 2.9%
claude-searchbot · 44.0% blocked 1.2%
claude-user · 37.9% blocked 1.1%
applebot-extended directive · 90.1% blocked 11.7%
applebot · 9.7% blocked 6.7%
bytespider · 68.1% blocked 17.4%
tiktokspider · 78.0% blocked 0.5%
cohere-ai · 69.6% blocked 2.7%
cohere-training-data-crawler · 83.2% blocked 0.7%
ccbot · 67.3% blocked 18.6%
deepseekbot · 80.4% blocked 0.3%
diffbot · 75.9% blocked 1.9%
duckassistbot · 58.6% blocked 1.0%
google-extended directive · 83.8% blocked 13.1%
googleother · 38.9% blocked 0.9%
petalbot · 34.0% blocked 8.0%
img2dataset · 68.3% blocked 1.0%
kangaroo bot · 63.5% blocked 0.6%
meta-externalagent · 86.9% blocked 12.3%
facebookbot · 64.2% blocked 2.3%
meta-externalfetcher · 56.0% blocked 1.1%
mistralai-user · 64.2% blocked 0.6%
gptbot · 81.6% blocked 15.5%
chatgpt-user · 52.0% blocked 4.8%
oai-searchbot · 29.9% blocked 3.0%
panscient.com · 95.3% blocked 0.2%
perplexitybot · 45.0% blocked 4.5%
perplexity-user · 36.8% blocked 1.2%
scrapy · 78.0% blocked 1.4%
timpibot · 77.6% blocked 1.4%
imgproxy · 98.4% blocked 0.1%
omgili · 84.2% blocked 2.2%
omgilibot · 83.0% blocked 2.2%
webzio-extended · 73.4% blocked 1.0%
youbot · 69.0% blocked 2.3%

Prior-art anchor (HTTP Almanac 2025, broad-web home-page census): GPTBot named on ~4.5% of files, up from ~2.9% the prior year — a near-doubling. Our top-1M numbers skew higher because popular sites have more (and more complex) robots.txt.