robots.txt & llms.txt across the Tranco top 1M

nytimes.com

Try: openai.com · nytimes.com · github.com · shopify.com

This site serves a robots.txt and blocks 25 of 26 named AI crawlers.

Tranco rank 153 Google posture open Crawls observed 2

AI crawlers

Every AI crawler this robots.txt names, grouped by who operates it, with the rule that actually applies under RFC 9309 group matching. Operators with the most blocked bots come first.

Anthropic 5 of 5 blocked
  • blocked Claude-SearchBot AI search index
  • blocked Claude-User live assistant fetch
  • blocked Claude-Web live assistant fetch
  • blocked ClaudeBot model training
  • blocked anthropic-ai undocumented
Meta 3 of 3 blocked
  • blocked FacebookBot model training
  • blocked meta-externalagent model training
  • blocked meta-externalfetcher live assistant fetch
OpenAI 3 of 3 blocked
  • blocked ChatGPT-User live assistant fetch
  • blocked GPTBot model training
  • blocked OAI-SearchBot AI search index
Perplexity 2 of 2 blocked
  • blocked Perplexity-User live assistant fetch
  • blocked PerplexityBot AI search index
Webz.io 2 of 2 blocked
  • blocked omgili model training
  • blocked omgilibot model training
Apple 1 of 1 blocked
  • blocked Applebot-Extended opt-out directive directive
ByteDance 1 of 1 blocked
  • blocked Bytespider model training
Cohere 1 of 1 blocked
  • blocked cohere-ai model training
Common Crawl 1 of 1 blocked
  • blocked CCBot model training
Diffbot 1 of 1 blocked
  • blocked Diffbot model training
DuckDuckGo 1 of 1 blocked
  • blocked DuckAssistBot live assistant fetch
Google 1 of 1 blocked
  • blocked Google-Extended opt-out directive directive
Scrapy 1 of 1 blocked
  • blocked Scrapy undocumented
Timpi 1 of 1 blocked
  • blocked Timpibot model training
You.com 1 of 1 blocked
  • blocked YouBot AI search index
Amazon 0 of 1 blocked
  • partly blocked AmazonBot live assistant fetch
15 other tracked AI crawlers are not named individually. They are partly restricted — they inherit the wildcard (*) group's rules.

robots.txt health

Fetch HTTP 200 the file was retrieved and parsed
Size 8,259 bytes within the 500 KiB limit crawlers must honour
Wildcard group present a User-agent: * group governs any crawler without its own rules
Sitemap 25 declared declaring a sitemap helps crawlers discover pages

llms.txt

CrawlVariantStatusConformant SectionsLinksGenerator
CC-MAIN-2026-30 404
CC-MAIN-2026-25 404

Crawl posture history

Google's interpretation of this robots.txt over status history (§3.3).

open CC-MAIN-2026-30 2xx: robots.txt present, no full disallow in effect
open CC-MAIN-2026-25 2xx: robots.txt present, no full disallow in effect

Live files on nytimes.com

The file as it exists on the host right now — may differ from what we measured. We link out rather than mirror (we store derived metrics, not raw bodies).

Raw observations

2 crawls, every source

A domain usually has more than one observation per crawl because we see it in the Common Crawl shard stream and fetch it ourselves. These are not duplicates — they are independent observations of the same file, and one is marked canonical. The rest of this page reads the canonical row.

CC-MAIN-2026-30

SourceRoleStatusSizeGroups AllowDisallowSitemaps
cc_shard
from the Common Crawl bulk shard stream
canonical 200 8.1 KiB 57 27 144 25

CC-MAIN-2026-25

SourceRoleStatusSizeGroups AllowDisallowSitemaps
self_crawl
fetched directly by our own polite crawler
canonical 200 8.1 KiB 57 27 144 25