robots.txt & llms.txt across the Tranco top 1M
Data as of 2026-09-07 · robots.txt via Common Crawl CC-MAIN-2026-34 + our polite crawl · llms.txt via our polite crawl · Tranco list XN67N (2026-09-03T22:00:02.533886) · 1,000,000 panel domains · crawls: 2026-09-07 · 2026-08-02 · 2026-07-25 · 2026-07-17 1,212,578 robots.txt fetches · 740,113 parsed · 56,498 llms.txt probes

indeed.com

This site serves a robots.txt and blocks 5 of 31 named AI crawlers.

Tranco rank 361 Observed snapshot 2026-09-07 Google posture open Crawls observed 2

How this site compares

Each row is one measured trait, not a composite score. Rank peers are Ranks 1–1,000; the second comparison is the full observed Tranco panel. Populations differ by metric and are printed with every rate.

TraitThis site Ranks 1–1,000Observed Tranco panel
robots.txt served Yes 53.3% of 1,000 61.7% of 1,212,578
robots.txt returned 404 No 4.6% of 1,000 8.4% of 1,212,578
Blanket disallow No 8.9% of 529 10.3% of 740,113
Effectively open file No 11.3% of 529 25.7% of 740,113
Wildcard group Yes 95.1% of 529 89.8% of 740,113
Sitemap declared No 66.0% of 529 60.7% of 740,113
Over the RFC parse limit No 0.8% of 529 0.3% of 740,113
How these comparisons are measured

Status rates use canonical robots.txt observations. File traits use only parsed canonical files. A missing metric stays unavailable rather than becoming 0%. Rank bands are exclusive, so Ranks 1–1,000 does not include more popular bands.

Treatment by purpose

Declared purpose comes from the versioned agent registry. These are separate policy dimensions, not inputs to a score. An unnamed purpose is not automatically allowed; it inherits the wildcard rule described below.

Model training Mixed
AI2Bot · Bytespider · CCBot · ClaudeBot · cohere-training-data-crawler · DeepSeekBot · Diffbot · FacebookBot · GPTBot · img2dataset · Meta-ExternalAgent · omgili · omgilibot · Timpibot · Webzio-Extended
AI search Mixed
Claude-SearchBot · OAI-SearchBot · PerplexityBot · PetalBot · YouBot
Live assistant fetch Mixed
AmazonBot · ChatGPT-User · Claude-User · DuckAssistBot · Meta-ExternalFetcher · Perplexity-User
Opt-out directives Mixed
Applebot-Extended · Google-Extended
Ambiguous / undocumented Mixed
anthropic-ai · GoogleOther · Scrapy

Named agents

Every tracked agent this robots.txt names, grouped by operator, with the rule that applies under RFC 9309 group matching. A peer targeting rate means sites that name the token; it is not a block rate.

Webz.io 3 of 3 blocked
  • blocked Webzio-Extended model training 7.8% of rank peers name it 65.6% block among 6,666 observed namers
  • blocked omgili model training 12.1% of rank peers name it 81.4% block among 16,162 observed namers
  • blocked omgilibot model training 9.1% of rank peers name it 80.4% block among 16,263 observed namers
Google 1 of 2 blocked
  • blocked GoogleOther undocumented 6.6% of rank peers name it 34.6% block among 7,220 observed namers
  • partly blocked Google-Extended opt-out directive directive 22.1% of rank peers name it 84.9% block among 108,283 observed namers
Scrapy 1 of 1 blocked
  • blocked Scrapy undocumented 9.5% of rank peers name it 73.3% block among 10,043 observed namers
Allen Institute 0 of 1 blocked
  • partly blocked AI2Bot model training 7.6% of rank peers name it 63.9% block among 9,650 observed namers
Amazon 0 of 1 blocked
  • partly blocked AmazonBot live assistant fetch 13.0% of rank peers name it 90.1% block among 107,461 observed namers
Anthropic 0 of 4 blocked
  • partly blocked Claude-SearchBot AI search index 12.3% of rank peers name it 40.3% block among 10,183 observed namers
  • partly blocked Claude-User live assistant fetch 13.0% of rank peers name it 34.4% block among 8,980 observed namers
  • partly blocked ClaudeBot model training 23.3% of rank peers name it 83.4% block among 117,092 observed namers
  • partly blocked anthropic-ai undocumented 12.7% of rank peers name it 62.5% block among 27,783 observed namers
Apple 0 of 1 blocked
  • partly blocked Applebot-Extended opt-out directive directive 15.7% of rank peers name it 90.3% block among 98,492 observed namers
ByteDance 0 of 1 blocked
  • partly blocked Bytespider model training 19.7% of rank peers name it 63.9% block among 154,750 observed namers
Cohere 0 of 1 blocked
  • partly blocked cohere-training-data-crawler model training 7.2% of rank peers name it 82.9% block among 4,332 observed namers
Common Crawl 0 of 1 blocked
  • partly blocked CCBot model training 21.7% of rank peers name it 63.3% block among 162,331 observed namers
DeepSeek 0 of 1 blocked
  • partly blocked DeepSeekBot model training 5.7% of rank peers name it 79.7% block among 2,304 observed namers
Diffbot 0 of 1 blocked
  • partly blocked Diffbot model training 12.5% of rank peers name it 72.2% block among 14,397 observed namers
DuckDuckGo 0 of 1 blocked
  • partly blocked DuckAssistBot live assistant fetch 9.6% of rank peers name it 55.6% block among 7,406 observed namers
Huawei 0 of 1 blocked
  • partly blocked PetalBot AI search index 11.9% of rank peers name it 26.9% block among 72,787 observed namers
Meta 0 of 3 blocked
  • partly blocked FacebookBot model training 10.0% of rank peers name it 61.7% block among 16,667 observed namers
  • partly blocked Meta-ExternalAgent model training 17.8% of rank peers name it 87.6% block among 100,566 observed namers
  • partly blocked Meta-ExternalFetcher live assistant fetch 8.7% of rank peers name it 49.6% block among 7,913 observed namers
OpenAI 0 of 3 blocked
  • partly blocked ChatGPT-User live assistant fetch 17.8% of rank peers name it 49.5% block among 35,854 observed namers
  • partly blocked GPTBot model training 25.0% of rank peers name it 82.8% block among 125,713 observed namers
  • partly blocked OAI-SearchBot AI search index 16.3% of rank peers name it 26.9% block among 23,333 observed namers
Perplexity 0 of 2 blocked
  • partly blocked Perplexity-User live assistant fetch 11.2% of rank peers name it 33.1% block among 9,189 observed namers
  • partly blocked PerplexityBot AI search index 21.4% of rank peers name it 42.2% block among 33,828 observed namers
Timpi 0 of 1 blocked
  • partly blocked Timpibot model training 8.9% of rank peers name it 72.5% block among 9,957 observed namers
You.com 0 of 1 blocked
  • partly blocked YouBot AI search index 10.8% of rank peers name it 65.9% block among 16,910 observed namers
img2dataset 0 of 1 blocked
  • partly blocked img2dataset model training 6.0% of rank peers name it 62.1% block among 6,942 observed namers
10 other tracked AI crawlers are not named individually. They are partly restricted — they inherit the wildcard (*) group's rules.

File evidence

/robots.txt Served and parsed Observed in Botfy's published crawl.
/llms.txt Not probed A related variant does not establish base-file adoption.

robots.txt details

Fetch HTTP 200 the file was retrieved and parsed
Size 13,097 bytes within the 500 KiB limit crawlers must honour
Wildcard group present a User-agent: * group governs any crawler without its own rules
Sitemap none declared no Sitemap: directive is present

Crawl posture history

Google's interpretation of this robots.txt over status history (§3.3).

open CC-MAIN-2026-34 2xx: robots.txt present, no full disallow in effect
open CC-MAIN-2026-30 2xx: robots.txt present, no full disallow in effect

Compare indeed.com

Put this site's measured policy beside another site. Botfy keeps missing evidence visible and does not assign a score or winner.

Live files on indeed.com

The file as it exists on the host right now — may differ from what we measured. We link out rather than mirror (we store derived metrics, not raw bodies).

Raw observations

2 crawls, every source

A domain usually has more than one observation per crawl because we see it in the Common Crawl shard stream and fetch it ourselves. These are not duplicates — they are independent observations of the same file, and one is marked canonical. The rest of this page reads the canonical row.

CC-MAIN-2026-34

SourceRoleStatusSizeGroups AllowDisallowSitemaps
cc_shard
from the Common Crawl bulk shard stream
canonical 200 12.8 KiB 9 37 468 0

CC-MAIN-2026-30

SourceRoleStatusSizeGroups AllowDisallowSitemaps
cc_shard
from the Common Crawl bulk shard stream
canonical 200 12.2 KiB 9 7 468 0

Inferred tech hypothesis

Derived from robots.txt patterns, not observed directly — labelled as a hypothesis because it is inference, not measurement.

Cloudflare · 95% Cloudflare · 95%