robots.txt & llms.txt across the Tranco top 1M
Data as of 2026-07-25 · robots.txt via Common Crawl CC-MAIN-2026-25 + our polite crawl · llms.txt via our polite crawl · Tranco top-1,000,000 · crawls: 2026-07-25 · 2026-07-17 687,794 robots.txt fetches · 491,025 parsed · 54,000 llms.txt probes

The state of robots.txt & llms.txt

Headline crawler-policy metrics across the Tranco top 1,000,000, sourced primarily from Common Crawl. All numbers are measured; prior-art baselines (Almanac 2025) are shown for context.

robots.txt returns 200
72.1%
Almanac 2025: 84.9%
absent (404)
8.3%
Almanac 2025: 13.0%
disallow-all
2.7%
blocks every crawler
allow-all
25.7%
effectively fully open
Google posture: open
92.7%
effective crawl posture (§3.3)
declares a sitemap
63.7%
Sitemap: directive present

Effective crawl posture

Google's interpretation of robots.txt over status history (§3.3).

data table
postureshare
open92.7%
blocked7.2%
cached0.1%
unknown0.0%

open vs blocked vs cached vs unknown.

Published → conformant → read

We measure adoption & conformance — never consumption. The "~3% ever read" bar is Ahrefs prior art (gray), not ours.

data table
11.1%
published
(our data)
~3%
ever read by AI
(Ahrefs prior art)

Of published files, 88.7% conform to the llmstxt.org spec.

AI-crawler targeting leaderboard

Share of parsed robots.txt files that name each AI user-agent in a directive — a targeting rate (the bot is addressed), not an allow/deny split. directive marks opt-out tokens (e.g. Google-Extended) that are not real request-log crawlers.

gptbot 21.1%
oai-searchbot 4.9%
chatgpt-user 7.2%
claudebot 18.5%
claude-user 1.7%
anthropic-ai 5.4%
ccbot 18.3%
google-extended directive 16.6%

See all AI crawlers →

Sample this crawl: 687,794 robots.txt fetches, 491,025 parsed, 54,000 llms.txt probes.