robots.txt & llms.txt across the Tranco top 1M
Data as of 2026-08-02 · robots.txt via Common Crawl CC-MAIN-2026-30 + our polite crawl · llms.txt via our polite crawl · Tranco list L5684 (2026-07-14) · 1,000,000 panel domains · crawls: 2026-09-07 · 2026-08-02 · 2026-07-25 · 2026-07-17 1,000,000 robots.txt fetches · 608,834 parsed · 55,770 llms.txt probes

Which AI crawlers do websites block?

Share of parsed robots.txt files across the top 1M that name each AI user-agent in a directive — a targeting rate (the bot is addressed by a rule), not an allow-vs-deny split (the data records that a bot is addressed, not the decision). Bars are tinted by vendor.

Weighting: Raw Rank-weighted
directive tokens (Google-Extended, Applebot-Extended) are robots.txt opt-out directives, NOT crawlers that appear in request logs. We model declared opt-out separately from actual-traffic presence (PRD FR-17b).

Related: which sites let AI crawlers in while blocking Google — the inverse configuration, and rarer than it looks.

Do sites that publish llms.txt also block AI crawlers?

The flagship cross-tab: does a site's llms.txt (a welcome mat for LLMs) agree with its robots.txt AI-crawler rules? The mixed cell — publishes an llms.txt and blocks AI crawlers — is a genuine contradiction.

data table
quadrantdomains
welcoming (llms.txt + allows AI)785
mixed (llms.txt + blocks AI)603
blocking (blocks AI, no llms.txt)97097
indifferent (neither)48376

Domains with a parsed robots.txt this crawl, crossed with genuine llms.txt presence (a 200 whose body is a real document). The mixed cell is the headline contradiction.

Which AI crawlers are named most often?

Anthropic 49.1% OpenAI 43.3% Google 24.8% Perplexity 22.7% Apple 21.9% Meta 20.3%
ai2bot · 67.8% blocked 3.1%
ai2bot-dolma · 66.7% blocked 2.3%
semanticscholarbot · 92.2% blocked 0.1%
amazonbot · 89.6% blocked 13.1%
claudebot · 82.4% blocked 18.0%
anthropic-ai · 64.2% blocked 10.7%
claude-web · 65.9% blocked 9.8%
claude-user · 37.9% blocked 5.3%
claude-searchbot · 44.0% blocked 5.3%
applebot-extended directive · 90.1% blocked 13.3%
applebot · 9.7% blocked 8.6%
bytespider · 68.1% blocked 12.6%
tiktokspider · 78.0% blocked 0.8%
cohere-ai · 69.6% blocked 9.7%
cohere-training-data-crawler · 83.2% blocked 2.6%
ccbot · 67.3% blocked 17.6%
deepseekbot · 80.4% blocked 2.1%
diffbot · 75.9% blocked 5.5%
duckassistbot · 58.6% blocked 4.3%
google-extended directive · 83.8% blocked 22.3%
googleother · 38.9% blocked 2.5%
petalbot · 34.0% blocked 9.9%
img2dataset · 68.3% blocked 2.8%
kangaroo bot · 63.5% blocked 2.2%
meta-externalagent · 86.9% blocked 10.7%
facebookbot · 64.2% blocked 5.5%
meta-externalfetcher · 56.0% blocked 4.1%
mistralai-user · 64.2% blocked 2.7%
gptbot · 81.6% blocked 23.3%
chatgpt-user · 52.0% blocked 12.5%
oai-searchbot · 29.9% blocked 7.4%
panscient.com · 95.3% blocked 1.3%
perplexitybot · 45.0% blocked 17.9%
perplexity-user · 36.8% blocked 4.9%
scrapy · 78.0% blocked 8.5%
timpibot · 77.6% blocked 4.0%
imgproxy · 98.4% blocked 0.9%
omgili · 84.2% blocked 5.6%
omgilibot · 83.0% blocked 4.6%
webzio-extended · 73.4% blocked 2.9%
youbot · 69.0% blocked 4.3%

Prior-art anchor (HTTP Almanac 2025, broad-web home-page census): GPTBot named on ~4.5% of files, up from ~2.9% the prior year — a near-doubling. Our top-1M numbers skew higher because popular sites have more (and more complex) robots.txt.