robots.txt & llms.txt across the Tranco top 1M
Data as of 2026-07-17 · robots.txt via Common Crawl CC-MAIN-2026-25 + our polite crawl · llms.txt via our polite crawl · Tranco top-1,000,000 · crawls: 2026-07-25 · 2026-07-17 746,000 robots.txt fetches · 574,664 parsed · 54,000 llms.txt probes

The state of robots.txt & llms.txt

Headline crawler-policy metrics across the Tranco top 1,000,000, sourced primarily from Common Crawl. All numbers are measured; prior-art baselines (Almanac 2025) are shown for context.

robots.txt returns 200
68.3%
Almanac 2025: 84.9%
absent (404)
12.0%
Almanac 2025: 13.0%
disallow-all
4.5%
blocks every crawler
allow-all
25.5%
effectively fully open
Google posture: open
91.7%
effective crawl posture (§3.3)
declares a sitemap
59.5%
Sitemap: directive present

Effective crawl posture

Google's interpretation of robots.txt over status history (§3.3).

data table
postureshare
open91.7%
blocked8.3%
cached0.0%
unknown0.0%

open vs blocked vs cached vs unknown.

Published → conformant → read

We measure adoption & conformance — never consumption. The "~3% ever read" bar is Ahrefs prior art (gray), not ours.

data table
11.1%
published
(our data)
~3%
ever read by AI
(Ahrefs prior art)

Of published files, 88.7% conform to the llmstxt.org spec.

AI-crawler targeting leaderboard

Share of parsed robots.txt files that name each AI user-agent in a directive — a targeting rate (the bot is addressed), not an allow/deny split. directive marks opt-out tokens (e.g. Google-Extended) that are not real request-log crawlers.

gptbot 18.8%
oai-searchbot 4.2%
chatgpt-user 6.2%
claudebot 16.5%
claude-user 1.4%
anthropic-ai 4.6%
ccbot 16.2%
google-extended directive 14.7%

See all AI crawlers →

Sample this crawl: 746,000 robots.txt fetches, 574,664 parsed, 54,000 llms.txt probes.