nytimes.com
Try: openai.com · nytimes.com · github.com · shopify.com
This site serves a robots.txt and blocks 25 of 26 named AI crawlers.
AI crawlers
Every AI crawler this robots.txt names, grouped by who operates it, with the rule that actually applies under RFC 9309 group matching. Operators with the most blocked bots come first.
- blocked Claude-SearchBot AI search index
- blocked Claude-User live assistant fetch
- blocked Claude-Web live assistant fetch
- blocked ClaudeBot model training
- blocked anthropic-ai undocumented
- blocked FacebookBot model training
- blocked meta-externalagent model training
- blocked meta-externalfetcher live assistant fetch
- blocked ChatGPT-User live assistant fetch
- blocked GPTBot model training
- blocked OAI-SearchBot AI search index
- blocked Perplexity-User live assistant fetch
- blocked PerplexityBot AI search index
- blocked omgili model training
- blocked omgilibot model training
- blocked Applebot-Extended opt-out directive directive
- blocked Bytespider model training
- blocked cohere-ai model training
- blocked CCBot model training
- blocked Diffbot model training
- blocked DuckAssistBot live assistant fetch
- blocked Google-Extended opt-out directive directive
- blocked Scrapy undocumented
- blocked Timpibot model training
- blocked YouBot AI search index
- partly blocked AmazonBot live assistant fetch
robots.txt health
llms.txt
| Crawl | Variant | Status | Conformant | Sections | Links | Generator |
|---|---|---|---|---|---|---|
| CC-MAIN-2026-30 | — | 404 | — | — | — | — |
| CC-MAIN-2026-25 | — | 404 | — | — | — | — |
Crawl posture history
Google's interpretation of this robots.txt over status history (§3.3).
Live files on nytimes.com
The file as it exists on the host right now — may differ from what we measured. We link out rather than mirror (we store derived metrics, not raw bodies).
/robots.txt ↗ /llms.txt ↗ /llms-full.txt ↗ /ai.txt ↗
Raw observations
2 crawls, every source
A domain usually has more than one observation per crawl because we see it in the Common Crawl shard stream and fetch it ourselves. These are not duplicates — they are independent observations of the same file, and one is marked canonical. The rest of this page reads the canonical row.
CC-MAIN-2026-30
| Source | Role | Status | Size | Groups | Allow | Disallow | Sitemaps |
|---|---|---|---|---|---|---|---|
| cc_shard from the Common Crawl bulk shard stream |
canonical | 200 | 8.1 KiB | 57 | 27 | 144 | 25 |
CC-MAIN-2026-25
| Source | Role | Status | Size | Groups | Allow | Disallow | Sitemaps |
|---|---|---|---|---|---|---|---|
| self_crawl fetched directly by our own polite crawler |
canonical | 200 | 8.1 KiB | 57 | 27 | 144 | 25 |