robots.txt & llms.txt across the Tranco top 1M
Data as of 2026-09-07 · robots.txt via Common Crawl CC-MAIN-2026-34 + our polite crawl · llms.txt via our polite crawl · Tranco list XN67N (2026-09-03T22:00:02.533886) · 1,000,000 panel domains 1,212,578 robots.txt fetches · 740,113 parsed · 56,498 llms.txt probes

Two public profiles · no score or winner

chatgpt.com and workers.dev

chatgpt.com #76 workers.dev #75

Measured takeaway

chatgpt.com's robots.txt disallows every crawler; workers.dev's does not.

Measured in Botfy's 2026-09-07 publication.

At a glance

A difference appears only when both sites have comparable evidence. Unavailable or unresolved observations remain incomplete.

Evidencechatgpt.com workers.dev
robots.txt Served and parsed Served and parsed
Google crawl posture Blocked Open
Unnamed crawler fallback Blocked the wildcard (*) group disallows everything, so they are blocked without being named Unrestricted no wildcard (*) group exists, so no rule applies to them at all
Base /llms.txt Not probed Related file observed at /ai.txt; /llms.txt not established

Treatment by purpose

Only resolved blocked, allowed, or mixed states count as a policy agreement or difference. Unnamed agents still inherit each site's wildcard rules.

Declared purposechatgpt.com workers.dev
Model training Blocked Not explicitly addressed
AI search Blocked Not explicitly addressed
Live assistant fetch Blocked Not explicitly addressed
Opt-out directives Blocked Not explicitly addressed
Ambiguous / undocumented Blocked Not explicitly addressed

Named agents

Measured differences appear first. “Not explicitly addressed” is not the same as allowed; the site's wildcard fallback governs that crawler.

Agentchatgpt.com workers.dev
CCBot Common Crawl Blocked Not explicitly addressed
Bytespider ByteDance Blocked Not explicitly addressed
FacebookBot Meta Blocked Not explicitly addressed
Omgili Webz.io Blocked Not explicitly addressed
Omgilibot Webz.io Blocked Not explicitly addressed
img2dataset img2dataset Blocked Not explicitly addressed
PerplexityBot Perplexity Blocked Not explicitly addressed
Claude-Web Anthropic Blocked Not explicitly addressed
Google-Extended Google Blocked Not explicitly addressed
anthropic-ai Anthropic Blocked Not explicitly addressed

How each site compares with its own rank peers

The sites may belong to different rank bands. Each percentage keeps its own population and denominator; Botfy does not subtract rates from unlike cohorts.

Traitchatgpt.com workers.dev
Wildcard group Yes 95.1% Ranks 1–1,000 · n=529 No 95.1% Ranks 1–1,000 · n=529
Effectively open file No 11.3% Ranks 1–1,000 · n=529 No 11.3% Ranks 1–1,000 · n=529
Sitemap declared Yes 66.0% Ranks 1–1,000 · n=529 No 66.0% Ranks 1–1,000 · n=529
robots.txt served Yes 53.3% Ranks 1–1,000 · n=1,000 Yes 53.3% Ranks 1–1,000 · n=1,000
robots.txt returned 404 No 4.6% Ranks 1–1,000 · n=1,000 No 4.6% Ranks 1–1,000 · n=1,000
Over the RFC parse limit No 0.8% Ranks 1–1,000 · n=529 Yes 0.8% Ranks 1–1,000 · n=529
Blanket disallow Yes 8.9% Ranks 1–1,000 · n=529 No 8.9% Ranks 1–1,000 · n=529
How this comparison is measured

Both sites use Botfy's latest published observations. A missing fetch, unresolved agent verdict, and measured absence are different states. Peer percentages use each site's exclusive Tranco rank band and print their own denominator. Populations, sources, and limits →

Keep comparing

Each suggestion replaces one side using an editorial collection or current rank proximity. Neither means the sites have similar policies.