The state of robots.txt & llms.txt
Headline crawler-policy metrics across the Tranco top 1,000,000, sourced primarily from Common Crawl. All numbers are measured; prior-art baselines (Almanac 2025) are shown for context.
Effective crawl posture
Google's interpretation of robots.txt over status history (§3.3).
data table
| posture | share |
|---|---|
| open | 91.7% |
| blocked | 8.3% |
| cached | 0.0% |
| unknown | 0.0% |
open vs blocked vs cached vs unknown.
Published → conformant → read
We measure adoption & conformance — never consumption. The "~3% ever read" bar is Ahrefs prior art (gray), not ours.
data table
Of published files, 88.7% conform to the llmstxt.org spec.
AI-crawler targeting leaderboard
Share of parsed robots.txt files that name each AI user-agent in a directive — a targeting rate (the bot is addressed), not an allow/deny split. directive marks opt-out tokens (e.g. Google-Extended) that are not real request-log crawlers.
Sample this crawl: 746,000 robots.txt fetches, 574,664 parsed, 54,000 llms.txt probes.