chatgpt.com
This site disallows everything and blocks 10 of 10 named AI crawlers.
How this site compares
Each row is one measured trait, not a composite score. Rank peers are Ranks 1–1,000; the second comparison is the full observed Tranco panel. Populations differ by metric and are printed with every rate.
| Trait | This site | Ranks 1–1,000 | Observed Tranco panel |
|---|---|---|---|
| robots.txt served | Yes | 53.3% of 1,000 | 61.7% of 1,212,578 |
| robots.txt returned 404 | No | 4.6% of 1,000 | 8.4% of 1,212,578 |
| Blanket disallow | Yes | 8.9% of 529 | 10.3% of 740,113 |
| Effectively open file | No | 11.3% of 529 | 25.7% of 740,113 |
| Wildcard group | Yes | 95.1% of 529 | 89.8% of 740,113 |
| Sitemap declared | Yes | 66.0% of 529 | 60.7% of 740,113 |
| Over the RFC parse limit | No | 0.8% of 529 | 0.3% of 740,113 |
How these comparisons are measured
Status rates use canonical robots.txt observations. File traits use only parsed canonical files. A missing metric stays unavailable rather than becoming 0%. Rank bands are exclusive, so Ranks 1–1,000 does not include more popular bands.
Treatment by purpose
Declared purpose comes from the versioned agent registry. These are separate policy dimensions, not inputs to a score. An unnamed purpose is not automatically allowed; it inherits the wildcard rule described below.
Named agents
Every tracked agent this robots.txt names, grouped by operator, with the rule that applies under RFC 9309 group matching. A peer targeting rate means sites that name the token; it is not a block rate.
- blocked Claude-Web live assistant fetch 11.3% of rank peers name it 63.9% block among 21,751 observed namers
- blocked anthropic-ai undocumented 12.7% of rank peers name it 62.5% block among 27,783 observed namers
- blocked Omgili model training 12.1% of rank peers name it 81.4% block among 16,162 observed namers
- blocked Omgilibot model training 9.1% of rank peers name it 80.4% block among 16,263 observed namers
- blocked Bytespider model training 19.7% of rank peers name it 63.9% block among 154,750 observed namers
- blocked CCBot model training 21.7% of rank peers name it 63.3% block among 162,331 observed namers
- blocked Google-Extended opt-out directive directive 22.1% of rank peers name it 84.9% block among 108,283 observed namers
- blocked FacebookBot model training 10.0% of rank peers name it 61.7% block among 16,667 observed namers
- blocked PerplexityBot AI search index 21.4% of rank peers name it 42.2% block among 33,828 observed namers
- blocked img2dataset model training 6.0% of rank peers name it 62.1% block among 6,942 observed namers
File evidence
robots.txt details
Crawl posture history
Google's interpretation of this robots.txt over status history (§3.3).
Compare chatgpt.com
Put this site's measured policy beside another site. Botfy keeps missing evidence visible and does not assign a score or winner.
Live files on chatgpt.com
The file as it exists on the host right now — may differ from what we measured. We link out rather than mirror (we store derived metrics, not raw bodies).
/robots.txt ↗ /llms.txt ↗ /llms-full.txt ↗ /ai.txt ↗
Raw observations
2 crawls, every source
A domain usually has more than one observation per crawl because we see it in the Common Crawl shard stream and fetch it ourselves. These are not duplicates — they are independent observations of the same file, and one is marked canonical. The rest of this page reads the canonical row.
CC-MAIN-2026-34
| Source | Role | Status | Size | Groups | Allow | Disallow | Sitemaps |
|---|---|---|---|---|---|---|---|
| cc_shard from the Common Crawl bulk shard stream |
canonical | 200 | 4.2 KiB | 14 | 156 | 20 | 6 |
CC-MAIN-2026-30
| Source | Role | Status | Size | Groups | Allow | Disallow | Sitemaps |
|---|---|---|---|---|---|---|---|
| cc_shard from the Common Crawl bulk shard stream |
canonical | 200 | 3.9 KiB | 14 | 144 | 20 | 5 |