robots.txt & llms.txt across the Tranco top 1M

Domain explorer

Look up a single domain's robots.txt & llms.txt metrics, effective AI policy, inferred tech stack (labelled hypothesis), and crawl history. Read-only; one indexed lookup per query.

facebook.com

Tranco rank
5
popularity rank
TLD
com
facebook.com
robots.txt records
1
across crawls
llms.txt records
1
across crawls

Live files on facebook.com

Open the file as it exists on the host right now — may differ from what we measured. We link out rather than mirror (we store derived metrics, not raw bodies).

Effective AI / crawl policy (Google posture, §3.3)

CrawlPostureConsec. 5xxReasonAs of
CC-MAIN-2026-25 blocked 0 2xx: robots.txt disallows all crawling (Disallow: /)

robots.txt

CrawlStatusBytesGroupsAllowDisallow SitemapsWildcard *Disallow-allAI tokensAI verdicts
CC-MAIN-2026-25 200 20,484 34 76 682 0 yes yes Amazonbot, Applebot-Extended, ClaudeBot, Google-Extended, GPTBot, PerplexityBot, PetalBot, Scrapy, Applebot amazonbot:block, applebot-extended:block, claudebot:block, google-extended:block, gptbot:block, perplexitybot:block, petalbot:block, scrapy:block, applebot:partial

llms.txt

Collected by our own polite crawler, not Common Crawl — Common Crawl doesn't capture /llms.txt. See Methodology.

CrawlStatusVariantConformantSections LinksGenerator
CC-MAIN-2026-25 200 /ai.txt no 0 0 Shopify

llms.txt data reflects adoption + conformance only, never consumption — publishing a file does not mean any AI reads it.

Inferred tech stack hypothesis

No tech inference for this domain.