robots.txt & llms.txt across the Tranco top 1M

github.com

Try: openai.com · nytimes.com · github.com · shopify.com

This site serves a robots.txt and names no AI crawler.

Tranco rank 30 Google posture open Crawls observed 2

AI crawlers

This robots.txt names no AI crawler.

41 other tracked AI crawlers are not named individually. They are partly restricted — they inherit the wildcard (*) group's rules.

robots.txt health

Fetch HTTP 200 the file was retrieved and parsed
Size 2,274 bytes within the 500 KiB limit crawlers must honour
Wildcard group present a User-agent: * group governs any crawler without its own rules
Sitemap none declared no Sitemap: directive is present

llms.txt

CrawlVariantStatusConformant SectionsLinksGenerator
CC-MAIN-2026-30 /llms.txt 200 yes 12 117
CC-MAIN-2026-25 /llms.txt 200 yes 12 117

Crawl posture history

Google's interpretation of this robots.txt over status history (§3.3).

open CC-MAIN-2026-30 2xx: robots.txt present, no full disallow in effect
open CC-MAIN-2026-25 2xx: robots.txt present, no full disallow in effect

Live files on github.com

The file as it exists on the host right now — may differ from what we measured. We link out rather than mirror (we store derived metrics, not raw bodies).

Raw observations

2 crawls, every source

A domain usually has more than one observation per crawl because we see it in the Common Crawl shard stream and fetch it ourselves. These are not duplicates — they are independent observations of the same file, and one is marked canonical. The rest of this page reads the canonical row.

CC-MAIN-2026-30

SourceRoleStatusSizeGroups AllowDisallowSitemaps
cc_shard
from the Common Crawl bulk shard stream
canonical 200 2.2 KiB 5 1 82 0

CC-MAIN-2026-25

SourceRoleStatusSizeGroups AllowDisallowSitemaps
cc_columnar
canonical 200 2.2 KiB 5 1 82 0