Which AI crawlers do websites block?
Share of parsed robots.txt files across the top 1M that name each AI user-agent in a directive — a targeting rate (the bot is addressed by a rule), not an allow-vs-deny split (the data records that a bot is addressed, not the decision). Bars are tinted by vendor.
Related: which sites let AI crawlers in while blocking Google — the inverse configuration, and rarer than it looks.
Do sites that publish llms.txt also block AI crawlers?
The flagship cross-tab: does a site's llms.txt (a welcome mat for LLMs) agree with its robots.txt AI-crawler rules? The mixed cell — publishes an llms.txt and blocks AI crawlers — is a genuine contradiction.
data table
| quadrant | domains |
|---|---|
| welcoming (llms.txt + allows AI) | 631 |
| mixed (llms.txt + blocks AI) | 560 |
| blocking (blocks AI, no llms.txt) | 56276 |
| indifferent (neither) | 14862 |
Domains with a parsed robots.txt this crawl, crossed with genuine llms.txt presence (a 200 whose body is a real document). The mixed cell is the headline contradiction.
Which AI crawlers are named most often?
Prior-art anchor (HTTP Almanac 2025, broad-web home-page census): GPTBot named on ~4.5% of files, up from ~2.9% the prior year — a near-doubling. Our top-1M numbers skew higher because popular sites have more (and more complex) robots.txt.