robots.txt & llms.txt across the Tranco top 1M
Data as of 2026-09-07 · robots.txt via Common Crawl CC-MAIN-2026-34 + our polite crawl · llms.txt via our polite crawl · Tranco list XN67N (2026-09-03T22:00:02.533886) · 1,000,000 panel domains · crawls: 2026-09-07 · 2026-08-02 · 2026-07-25 · 2026-07-17 1,212,578 robots.txt fetches · 740,113 parsed · 56,498 llms.txt probes

x.com

This site disallows everything and names 4 AI crawlers, blocking none.

Tranco rank 50 Observed snapshot 2026-09-07 Google posture blocked Crawls observed 2

How this site compares

Each row is one measured trait, not a composite score. Rank peers are Ranks 1–1,000; the second comparison is the full observed Tranco panel. Populations differ by metric and are printed with every rate.

TraitThis site Ranks 1–1,000Observed Tranco panel
robots.txt served Yes 53.3% of 1,000 61.7% of 1,212,578
robots.txt returned 404 No 4.6% of 1,000 8.4% of 1,212,578
Blanket disallow Yes 8.9% of 529 10.3% of 740,113
Effectively open file No 11.3% of 529 25.7% of 740,113
Wildcard group Yes 95.1% of 529 89.8% of 740,113
Sitemap declared Yes 66.0% of 529 60.7% of 740,113
Over the RFC parse limit No 0.8% of 529 0.3% of 740,113
How these comparisons are measured

Status rates use canonical robots.txt observations. File traits use only parsed canonical files. A missing metric stays unavailable rather than becoming 0%. Rank bands are exclusive, so Ranks 1–1,000 does not include more popular bands.

Treatment by purpose

Declared purpose comes from the versioned agent registry. These are separate policy dimensions, not inputs to a score. An unnamed purpose is not automatically allowed; it inherits the wildcard rule described below.

Model training Mixed
FacebookBot · meta-externalagent
AI search Not explicitly addressed
No measured named token in this purpose.
Live assistant fetch Mixed
meta-externalfetcher
Opt-out directives Mixed
Google-Extended
Ambiguous / undocumented Not explicitly addressed
No measured named token in this purpose.

Named agents

Every tracked agent this robots.txt names, grouped by operator, with the rule that applies under RFC 9309 group matching. A peer targeting rate means sites that name the token; it is not a block rate.

Google 0 of 1 blocked
  • partly blocked Google-Extended opt-out directive directive 22.1% of rank peers name it 84.9% block among 108,283 observed namers
Meta 0 of 3 blocked
  • partly blocked FacebookBot model training 10.0% of rank peers name it 61.7% block among 16,667 observed namers
  • partly blocked meta-externalagent model training 17.8% of rank peers name it 87.6% block among 100,566 observed namers
  • partly blocked meta-externalfetcher live assistant fetch 8.7% of rank peers name it 49.6% block among 7,913 observed namers
37 other tracked AI crawlers are not named individually. They are blocked — the wildcard (*) group disallows everything, so they are blocked without being named.

File evidence

/robots.txt Served and parsed Observed in Botfy's published crawl.
/llms.txt Related file observed at /ai.txt; /llms.txt not established A related variant does not establish base-file adoption.

robots.txt details

Fetch HTTP 200 the file was retrieved and parsed
Size 2,678 bytes within the 500 KiB limit crawlers must honour
Wildcard group present a User-agent: * group governs any crawler without its own rules
Sitemap 1 declared declaring a sitemap helps crawlers discover pages

llms.txt observations

CrawlVariantStatusConformant SectionsLinksGenerator
CC-MAIN-2026-34 /ai.txt 200 no 0 6 Shopify
CC-MAIN-2026-30 /ai.txt 200 no 0 6 Shopify

Crawl posture history

Google's interpretation of this robots.txt over status history (§3.3).

blocked CC-MAIN-2026-34 2xx: robots.txt disallows all crawling (Disallow: /)
blocked CC-MAIN-2026-30 2xx: robots.txt disallows all crawling (Disallow: /)

Compare x.com

Put this site's measured policy beside another site. Botfy keeps missing evidence visible and does not assign a score or winner.

Live files on x.com

The file as it exists on the host right now — may differ from what we measured. We link out rather than mirror (we store derived metrics, not raw bodies).

Raw observations

2 crawls, every source

A domain usually has more than one observation per crawl because we see it in the Common Crawl shard stream and fetch it ourselves. These are not duplicates — they are independent observations of the same file, and one is marked canonical. The rest of this page reads the canonical row.

CC-MAIN-2026-34

SourceRoleStatusSizeGroups AllowDisallowSitemaps
cc_shard
from the Common Crawl bulk shard stream
canonical 200 2.6 KiB 10 17 45 1

CC-MAIN-2026-30

SourceRoleStatusSizeGroups AllowDisallowSitemaps
cc_shard
from the Common Crawl bulk shard stream
canonical 200 2.9 KiB 7 24 61 1