openai.com
Try: openai.com · nytimes.com · github.com · shopify.com
This site serves a robots.txt and names no AI crawler.
AI crawlers
This robots.txt names no AI crawler.
robots.txt health
Crawl posture history
Google's interpretation of this robots.txt over status history (§3.3).
Live files on openai.com
The file as it exists on the host right now — may differ from what we measured. We link out rather than mirror (we store derived metrics, not raw bodies).
/robots.txt ↗ /llms.txt ↗ /llms-full.txt ↗ /ai.txt ↗
Raw observations
2 crawls, every source
A domain usually has more than one observation per crawl because we see it in the Common Crawl shard stream and fetch it ourselves. These are not duplicates — they are independent observations of the same file, and one is marked canonical. The rest of this page reads the canonical row.
CC-MAIN-2026-30
| Source | Role | Status | Size | Groups | Allow | Disallow | Sitemaps |
|---|---|---|---|---|---|---|---|
| cc_shard from the Common Crawl bulk shard stream |
canonical | 200 | 98 B | 1 | 1 | 1 | 1 |
CC-MAIN-2026-25
| Source | Role | Status | Size | Groups | Allow | Disallow | Sitemaps |
|---|---|---|---|---|---|---|---|
| cc_columnar |
canonical | 200 | 98 B | 1 | 1 | 1 | 1 |