Domain explorer
Look up a single domain's robots.txt & llms.txt metrics, effective AI policy, inferred tech stack (labelled hypothesis), and crawl history. Read-only; one indexed lookup per query.
amazon.com
Tranco rank
25
popularity rank
TLD
com
amazon.com
robots.txt records
2
across crawls
llms.txt records
1
across crawls
Live files on amazon.com
Open the file as it exists on the host right now — may differ from what we measured. We link out rather than mirror (we store derived metrics, not raw bodies).
Effective AI / crawl policy (Google posture, §3.3)
| Crawl | Posture | Consec. 5xx | Reason | As of |
|---|---|---|---|---|
| CC-MAIN-2026-25 | open | 0 | 2xx: robots.txt present, no full disallow in effect | 2026-06-18T11:05:07-07:00 |
robots.txt
| Crawl | Status | Bytes | Groups | Allow | Disallow | Sitemaps | Wildcard * | Disallow-all | AI tokens | AI verdicts |
|---|---|---|---|---|---|---|---|---|---|---|
| CC-MAIN-2026-25 | 200 | 7,887 | 101 | 17 | 218 | 0 | yes | no | GPTBot, CCBot, PerplexityBot, Google-Extended, ClaudeBot, meta-externalagent, Bytespider, Scrapy, PetalBot, omgili, AI2Bot, MistralAI-User, Diffbot, DuckAssistBot, Ai2Bot-Dolma, meta-externalfetcher, cohere-ai, img2dataset, YouBot, Claude-User, Claude-SearchBot, Perplexity-User, imgproxy, ChatGPT-User, OAI-SearchBot, Timpibot, Claude-Web, cohere-training-data-crawler, DeepSeekBot, GoogleOther, Kangaroo Bot, panscient.com, webzio-extended | gptbot:block, ccbot:block, perplexitybot:block, google-extended:block, claudebot:block, meta-externalagent:block, bytespider:block, scrapy:block, petalbot:block, omgili:block, ai2bot:block, mistralai-user:block, diffbot:block, duckassistbot:block, ai2bot-dolma:block, meta-externalfetcher:block, cohere-ai:block, img2dataset:block, youbot:block, claude-user:block, claude-searchbot:block, perplexity-user:block, imgproxy:block, chatgpt-user:block, oai-searchbot:block, timpibot:block, claude-web:block, cohere-training-data-crawler:block, deepseekbot:block, googleother:block, kangaroo bot:block, panscient.com:block, webzio-extended:block |
| CC-MAIN-2026-25 | 200 | 7,887 | 101 | 17 | 218 | 0 | yes | no | GPTBot, CCBot, PerplexityBot, Google-Extended, ClaudeBot, meta-externalagent, Bytespider, Scrapy, PetalBot, omgili, AI2Bot, MistralAI-User, Diffbot, DuckAssistBot, Ai2Bot-Dolma, meta-externalfetcher, cohere-ai, img2dataset, YouBot, Claude-User, Claude-SearchBot, Perplexity-User, imgproxy, ChatGPT-User, OAI-SearchBot, Timpibot, Claude-Web, cohere-training-data-crawler, DeepSeekBot, GoogleOther, Kangaroo Bot, panscient.com, webzio-extended | gptbot:block, ccbot:block, perplexitybot:block, google-extended:block, claudebot:block, meta-externalagent:block, bytespider:block, scrapy:block, petalbot:block, omgili:block, ai2bot:block, mistralai-user:block, diffbot:block, duckassistbot:block, ai2bot-dolma:block, meta-externalfetcher:block, cohere-ai:block, img2dataset:block, youbot:block, claude-user:block, claude-searchbot:block, perplexity-user:block, imgproxy:block, chatgpt-user:block, oai-searchbot:block, timpibot:block, claude-web:block, cohere-training-data-crawler:block, deepseekbot:block, googleother:block, kangaroo bot:block, panscient.com:block, webzio-extended:block |
llms.txt
Collected by our own polite crawler, not Common Crawl — Common Crawl doesn't capture /llms.txt. See Methodology.
| Crawl | Status | Variant | Conformant | Sections | Links | Generator |
|---|---|---|---|---|---|---|
| CC-MAIN-2026-25 | 404 | — | — | — | — | — |
llms.txt data reflects adoption + conformance only, never consumption — publishing a file does not mean any AI reads it.
Inferred tech stack hypothesis
No tech inference for this domain.