AI crawlers
Share of parsed robots.txt files across the top 1M that name each AI user-agent in a directive — a targeting rate (the bot is addressed by a rule), not an allow-vs-deny split (the data records that a bot is addressed, not the decision). Bars are tinted by vendor.
Mixed signals: llms.txt welcome mat vs robots.txt block
The flagship cross-tab: does a site's llms.txt (a welcome mat for LLMs) agree with its robots.txt AI-crawler rules? The mixed cell — publishes an llms.txt and blocks AI crawlers — is a genuine contradiction.
data table
| quadrant | domains |
|---|---|
| welcoming (llms.txt + allows AI) | 631 |
| mixed (llms.txt + blocks AI) | 560 |
| blocking (blocks AI, no llms.txt) | 56276 |
| indifferent (neither) | 14862 |
Domains with a parsed robots.txt this crawl, crossed with genuine (200, non-HTML) llms.txt presence. The mixed cell is the headline contradiction.
Targeting leaderboard
Prior-art anchor (HTTP Almanac 2025, broad-web home-page census): GPTBot named on ~4.5% of files, up from ~2.9% the prior year — a near-doubling. Our top-1M numbers skew higher because popular sites have more (and more complex) robots.txt.