robots.txt & llms.txt across the Tranco top 1M
Data as of 2026-09-07 · robots.txt via Common Crawl CC-MAIN-2026-34 + our polite crawl · llms.txt via our polite crawl · Tranco list XN67N (2026-09-03T22:00:02.533886) · 1,000,000 panel domains · crawls: 2026-09-07 · 2026-08-02 · 2026-07-25 · 2026-07-17 1,212,578 robots.txt fetches · 740,113 parsed · 56,498 llms.txt probes

pvp.net

No parsable robots.txt — the URL did not respond.

Tranco rank 674 Observed snapshot 2026-09-07 Google posture blocked Crawls observed 2

How this site compares

Each row is one measured trait, not a composite score. Rank peers are Ranks 1–1,000; the second comparison is the full observed Tranco panel. Populations differ by metric and are printed with every rate.

TraitThis site Ranks 1–1,000Observed Tranco panel
robots.txt served Fetch failed: dns_error 53.3% of 1,000 61.7% of 1,212,578
robots.txt returned 404 Fetch failed: dns_error 4.6% of 1,000 8.4% of 1,212,578
Blanket disallow Parsing unavailable 8.9% of 529 10.3% of 740,113
Effectively open file Parsing unavailable 11.3% of 529 25.7% of 740,113
Wildcard group Parsing unavailable 95.1% of 529 89.8% of 740,113
Sitemap declared Parsing unavailable 66.0% of 529 60.7% of 740,113
Over the RFC parse limit Parsing unavailable 0.8% of 529 0.3% of 740,113
How these comparisons are measured

Status rates use canonical robots.txt observations. File traits use only parsed canonical files. A missing metric stays unavailable rather than becoming 0%. Rank bands are exclusive, so Ranks 1–1,000 does not include more popular bands.

Treatment by purpose

Declared purpose comes from the versioned agent registry. These are separate policy dimensions, not inputs to a score. An unnamed purpose is not automatically allowed; it inherits the wildcard rule described below.

Model training Observation unavailable
No measured named token in this purpose.
AI search Observation unavailable
No measured named token in this purpose.
Live assistant fetch Observation unavailable
No measured named token in this purpose.
Opt-out directives Observation unavailable
No measured named token in this purpose.
Ambiguous / undocumented Observation unavailable
No measured named token in this purpose.

Named agents

File evidence

/robots.txt Fetch failed: dns_error Observed in Botfy's published crawl.
/llms.txt Not probed A related variant does not establish base-file adoption.

robots.txt details

Fetch dns_error no parsable robots.txt was served at this URL

Crawl posture history

Google's interpretation of this robots.txt over status history (§3.3).

blocked CC-MAIN-2026-34 5xx/429/unreachable within 30d, no valid cached copy → treated as fully disallowed
blocked CC-MAIN-2026-30 5xx/429/unreachable within 30d, no valid cached copy → treated as fully disallowed

Compare pvp.net

Put this site's measured policy beside another site. Botfy keeps missing evidence visible and does not assign a score or winner.

Live files on pvp.net

The file as it exists on the host right now — may differ from what we measured. We link out rather than mirror (we store derived metrics, not raw bodies).

Raw observations

2 crawls, every source

A domain usually has more than one observation per crawl because we see it in the Common Crawl shard stream and fetch it ourselves. These are not duplicates — they are independent observations of the same file, and one is marked canonical. The rest of this page reads the canonical row.

CC-MAIN-2026-34

SourceRoleStatusSizeGroups AllowDisallowSitemaps
self_crawl
fetched directly by our own polite crawler
canonical dns_error

CC-MAIN-2026-30

SourceRoleStatusSizeGroups AllowDisallowSitemaps
self_crawl
fetched directly by our own polite crawler
canonical connect_error