robots.txt & llms.txt across the Tranco top 1M
Data as of 2026-07-25 · robots.txt via Common Crawl CC-MAIN-2026-25 + our polite crawl · llms.txt via our polite crawl · Tranco list ID not recorded for this snapshot · crawls: 2026-09-07 · 2026-08-02 · 2026-07-25 · 2026-07-17 687,794 robots.txt fetches · 491,025 parsed · 54,000 llms.txt probes

Which sites let AI crawlers in but block Google?

This snapshot cannot answer that question yet.

The cross-tab needs resolved allow/deny verdicts for each named AI crawler, and the 2026-07-25 snapshot does not carry them for enough domains to report. It is not that no site does this — it is that this crawl cannot tell you. The figure appears here once a crawl with per-bot verdicts is published.

How this was measured

Google posture is the effective posture under Google's documented behaviour for status codes, redirects and persistent errors, not merely what the file says.

"Allows an AI crawler" means the resolved verdict for a registry crawler is allow under RFC 9309 group matching. Both figures are measured on domains with a parsed HTTP 200 robots.txt in this crawl. Full definitions on the methodology page.