robots.txt & llms.txt across the Tranco top 1M
Data as of 2026-08-02 · robots.txt via Common Crawl CC-MAIN-2026-30 + our polite crawl · llms.txt via our polite crawl · Tranco list L5684 (2026-07-14) · 1,000,000 panel domains · crawls: 2026-09-07 · 2026-08-02 · 2026-07-25 · 2026-07-17 1,000,000 robots.txt fetches · 608,834 parsed · 55,770 llms.txt probes

Which sites let AI crawlers in but block Google?

1,345 domains in the Tranco top 1,000,000 shut Googlebot out while explicitly allowing at least one AI crawler, measured on 2026-08-02. That is the number worth quoting — and it is 31,376 before the deflation below.

Why the raw number is 31,376 and the real one is 1,345

The obvious query — Google blocked, an AI crawler explicitly allowed — returns 31,376 domains. 30,031 of them (96%) serve one of a small number of byte-identical files. Same group count, same row count, same disallow count, same allow-set, across tens of thousands of hosts. Those are generated networks, not crawler-policy decisions, and counting them would build a headline almost entirely out of spam infrastructure.

The exclusion is a stated mechanical rule, not a judgement about which sites look disreputable: a file signature shared by at least 1,000 distinct hosts is one template, counted once. A rule can be argued with. It also generalises — the second cluster below was found by the rule, not by anyone noticing it.

HostsAllowsGroupsRowsExample hosts
28,277 applebot, bytespider, ccbot, petalbot 16 69 889992712.xyz, 889992719.xyz, 889992722.xyz
1,754 meta-externalagent 40 119 000betpk.com, 000p999.com, 000q789.com

Is this just the tail of the web?

No. The independent cohort reaches the top of the panel, which is what makes it a decision rather than an artefact of abandoned domains.

Rank bandDomains
Top 10k17
10k–100k91
100k–500k660
500k–1M577

Which sites are they?

The best-ranked 100 of the 1,345, each linking to its full measured profile. "Allows" lists only crawlers the file names with an explicit Allow — a bot the file never mentions is not counted as allowed, because under RFC 9309 §2.2.1 it falls to the * group, and that group is exactly what blocks Google here.

RankSiteExplicitly allows
257 linktr.ee applebot-extended, google-extended
753 adform.net googleother
1,806 npci.org.in anthropic-ai, applebot, applebot-extended, ccbot, chatgpt-user, claude-web, claudebot, facebookbot, google-extended, googleother, gptbot, meta-externalagent, oai-searchbot, perplexitybot
2,043 redditstatic.com applebot
2,286 ad.gt googleother
2,526 news.com.au applebot
2,594 kuaishou.com chatgpt-user, deepseekbot, gptbot, oai-searchbot, perplexity-user, perplexitybot
2,960 theregister.com applebot, chatgpt-user, claude-searchbot, claude-user, oai-searchbot, perplexitybot
3,778 sportradarserving.com googleother
4,720 enlightenment.org applebot
6,474 hackernoon.com applebot, chatgpt-user, claude-searchbot, claude-user, claude-web, duckassistbot, facebookbot, oai-searchbot, perplexitybot
6,540 savana.com applebot, oai-searchbot
7,069 impact-ad.jp googleother
7,099 theaustralian.com.au applebot, chatgpt-user, gptbot, oai-searchbot
7,615 cargocollective.com chatgpt-user
8,786 realtimetrains.co.uk applebot
9,957 kaixo.com applebot, gptbot
11,050 umt.edu amazonbot, applebot, ccbot, claude-searchbot, claude-user, claudebot, google-extended, gptbot, meta-externalagent, oai-searchbot, perplexitybot, youbot
11,614 ele.me bytespider
11,700 musinsa.com applebot, chatgpt-user, claude-searchbot, claude-user, oai-searchbot, perplexity-user
11,766 link.me applebot
11,892 yospace.com applebot
13,446 kbs.co.kr chatgpt-user, gptbot
13,710 dailytelegraph.com.au applebot
14,063 hupu.com bytespider
14,509 iteratehq.com amazonbot, applebot-extended, bytespider, ccbot, chatgpt-user, claude-searchbot, claudebot, google-extended, gptbot, meta-externalagent, perplexitybot
14,641 heraldsun.com.au applebot
18,178 couriermail.com.au applebot
18,294 vagaro.com applebot, chatgpt-user, oai-searchbot, perplexity-user, perplexitybot
19,306 saashr.com googleother
19,311 iqair.com amazonbot, applebot, chatgpt-user, claude-searchbot, claude-user, claudebot, google-extended, googleother, gptbot, meta-externalagent, oai-searchbot, perplexitybot
19,647 zuoyebang.com bytespider
20,029 linkme.global applebot
20,690 africa.com anthropic-ai, applebot, chatgpt-user, claudebot, google-extended, gptbot, oai-searchbot, perplexitybot
21,542 accedo.tv chatgpt-user, claude-searchbot, claude-user, duckassistbot, mistralai-user, oai-searchbot, perplexity-user, perplexitybot, youbot
21,679 fow.lol applebot
22,482 melon.com applebot, meta-externalagent, meta-externalfetcher
24,762 upichalega.com anthropic-ai, applebot, applebot-extended, ccbot, chatgpt-user, claude-web, claudebot, facebookbot, google-extended, googleother, gptbot, meta-externalagent, oai-searchbot, perplexitybot
26,751 aldi.us chatgpt-user
27,330 adelaidenow.com.au applebot
27,973 fotki.com applebot, bytespider, tiktokspider
28,211 mansionglobal.com applebot, chatgpt-user, googleother, gptbot, oai-searchbot, petalbot
28,865 chilipiper.com chatgpt-user, claude-searchbot, claude-user, claudebot, google-extended, gptbot, oai-searchbot, perplexitybot
29,792 abqjournal.com meta-externalfetcher
29,991 shahid.net applebot
30,490 taste.com.au applebot
34,430 zhangyue.com bytespider, petalbot
35,060 coroflot.com applebot, claude-searchbot, claude-user
35,513 rstyle.me bytespider, meta-externalfetcher
37,801 headwayapp.co applebot
40,455 diytrade.com applebot, gptbot
40,584 musicstore.de anthropic-ai, applebot-extended, chatgpt-user, claudebot, cohere-ai, google-extended, gptbot, perplexitybot, youbot
41,107 chrobinson.com ai2bot, amazonbot, anthropic-ai, applebot, applebot-extended, bytespider, ccbot, chatgpt-user, claude-web, claudebot, cohere-ai, diffbot, duckassistbot, facebookbot, gptbot, meta-externalagent, mistralai-user, oai-searchbot, omgili, perplexity-user, perplexitybot, timpibot, youbot
41,836 kuaishou.cn chatgpt-user, deepseekbot, gptbot, oai-searchbot, perplexity-user, perplexitybot
42,135 apamanshop.com applebot, ccbot, claudebot, google-extended, gptbot
43,071 csudh.edu applebot
44,966 huaren.us googleother
45,047 vogue.com.au applebot
45,359 o18.link facebookbot
48,270 chenzhongtech.com chatgpt-user, deepseekbot, gptbot, oai-searchbot, perplexity-user, perplexitybot
52,759 vezeeta.com chatgpt-user
52,942 samsungcard.com claudebot, google-extended, gptbot
53,428 o18.click facebookbot
53,660 entertimeonline.com googleother
54,766 referralcandy.com applebot
55,305 diario.mx amazonbot, applebot, chatgpt-user, facebookbot, googleother, petalbot
57,741 nabd.com meta-externalagent
58,284 epochtimes.de applebot
63,130 cke.gov.pl scrapy
63,205 dd373.com amazonbot, applebot, bytespider, chatgpt-user, gptbot
63,349 sumtotal.host applebot
64,056 arcade-museum.com applebot, claude-searchbot, claude-user, oai-searchbot, perplexitybot
64,225 nea.gov.sg chatgpt-user, claude-searchbot, claude-user, oai-searchbot, perplexitybot
65,365 flylevel.com applebot, google-extended
66,097 ald12345.com amazonbot, applebot, chatgpt-user, deepseekbot, google-extended, gptbot, perplexitybot, youbot
67,238 dbnl.org applebot
67,496 manager.ro applebot, ccbot, gptbot
72,077 omantel.om chatgpt-user, claude-searchbot, claude-user, claudebot, google-extended, googleother, gptbot, oai-searchbot
72,501 accredible.com anthropic-ai, applebot, applebot-extended, chatgpt-user, claude-searchbot, claude-user, claudebot, cohere-ai, google-extended, googleother, gptbot, meta-externalagent, meta-externalfetcher, oai-searchbot, perplexity-user, perplexitybot
72,698 eisamay.com amazonbot, applebot-extended, bytespider, ccbot, claude-searchbot, claudebot, google-extended, gptbot, meta-externalagent, oai-searchbot, perplexitybot
74,456 qad.com anthropic-ai, claudebot, cohere-ai, gptbot, perplexitybot
75,036 blbet.com meta-externalagent
75,348 mvs.gov.ua applebot
76,135 bildelsbasen.se chatgpt-user, claude-searchbot, claude-user, duckassistbot, facebookbot, oai-searchbot, perplexity-user, perplexitybot, youbot
77,159 hzhyccgs.com bytespider
78,474 musicstore.com anthropic-ai, applebot-extended, chatgpt-user, claudebot, cohere-ai, google-extended, gptbot, perplexitybot, youbot
80,919 spotonflorida.com applebot
82,879 delicious.com.au applebot
84,061 rupay.co.in anthropic-ai, applebot, applebot-extended, ccbot, chatgpt-user, claude-web, claudebot, facebookbot, google-extended, googleother, gptbot, meta-externalagent, oai-searchbot, perplexitybot
85,012 365pk365.com meta-externalagent
85,441 bodyandsoul.com.au applebot
85,602 dirtycode.io applebot
87,543 themercury.com.au applebot
89,875 yrno.cz applebot, facebookbot, google-extended, googleother, meta-externalagent
91,174 999bd.com meta-externalagent
92,261 coinone.co.kr applebot, chatgpt-user, claude-searchbot, claude-user, oai-searchbot, perplexity-user, perplexitybot
93,323 ticketsolve.com applebot, chatgpt-user, claude-searchbot, claude-user, claudebot, google-extended, oai-searchbot, perplexity-user, perplexitybot
93,867 663bet.co meta-externalagent
93,913 aladdin288.com amazonbot, applebot, chatgpt-user, deepseekbot, google-extended, gptbot, perplexitybot, youbot
93,955 ar999e.com meta-externalagent

What this means if you run a site

Blocking Google and allowing AI is a real configuration, and it is rare. 1,345 domains out of a million is roughly one in 743. If you find your own site here and did not intend it, the usual cause is a blanket Disallow: / in the wildcard group with a later, narrower group naming an AI crawler — the AI rule wins for that bot and the wildcard still blocks everything else, Googlebot included.

How this was measured

Google posture is the effective posture under Google's documented behaviour for status codes, redirects and persistent errors, not merely what the file says. Of the 31,376 raw matches, the overwhelming majority reach blocked through an explicit Disallow: / rather than through the 30-day degradation path that follows repeated 5xx responses — so this is deliberate configuration, not breakage.

"Allows an AI crawler" means the resolved verdict for a registry crawler is allow under RFC 9309 group matching. Both figures are measured on domains with a parsed HTTP 200 robots.txt in this crawl. Full definitions on the methodology page.