robots.txt & llms.txt across the Tranco top 1M
Data as of 2026-09-07 · robots.txt via Common Crawl CC-MAIN-2026-34 + our polite crawl · llms.txt via our polite crawl · Tranco list XN67N (2026-09-03T22:00:02.533886) · 1,000,000 panel domains 1,212,578 robots.txt fetches · 740,113 parsed · 56,498 llms.txt probes

Two public profiles · no score or winner

debate.com.mx and hoteltonight.com

debate.com.mx #38,568 hoteltonight.com #38,570

Measured takeaway

debate.com.mx has mixed rules for model-training crawlers; hoteltonight.com blocks model-training crawlers.

debate.com.mx partly blocks GPTBot; hoteltonight.com blocks GPTBot.

Measured in Botfy's 2026-09-07 publication.

At a glance

A difference appears only when both sites have comparable evidence. Unavailable or unresolved observations remain incomplete.

Evidencedebate.com.mx hoteltonight.com
robots.txt Served and parsed Served and parsed
Google crawl posture Blocked Open
Unnamed crawler fallback Blocked the wildcard (*) group disallows everything, so they are blocked without being named Partly Restricted they inherit the wildcard (*) group's rules
Base /llms.txt Not probed Related file observed at /ai.txt; /llms.txt not established

Treatment by purpose

Only resolved blocked, allowed, or mixed states count as a policy agreement or difference. Unnamed agents still inherit each site's wildcard rules.

Declared purposedebate.com.mx hoteltonight.com
Model training Mixed Blocked
AI search Mixed Mixed
Live assistant fetch Mixed Mixed
Opt-out directives Mixed Blocked
Ambiguous / undocumented Mixed Blocked

Named agents

Measured differences appear first. “Not explicitly addressed” is not the same as allowed; the site's wildcard fallback governs that crawler.

Agentdebate.com.mx hoteltonight.com
GPTBot OpenAI Partly blocked Blocked
ClaudeBot Anthropic Partly blocked Blocked
Meta-ExternalAgent Meta Partly blocked Blocked
Google-Extended Google Partly blocked Blocked
Applebot-Extended Apple Partly blocked Blocked
anthropic-ai Anthropic Partly blocked Blocked
OAI-Searchbot OpenAI Partly blocked Partly blocked
Claude-SearchBot Anthropic Partly blocked Partly blocked
PerplexityBot Perplexity Partly blocked Partly blocked
ChatGPT-User OpenAI Partly blocked Partly blocked
CCBot Common Crawl Not explicitly addressed Blocked
Bytespider ByteDance Not explicitly addressed Blocked
8 additional tracked agents
Agentdebate.com.mx hoteltonight.com
FacebookBot Partly blockedNot explicitly addressed
Applebot Partly blockedNot explicitly addressed
YouBot Not explicitly addressedPartly blocked
PetalBot Partly blockedNot explicitly addressed
Claude-User Not explicitly addressedPartly blocked
Perplexity-User Not explicitly addressedPartly blocked
DuckAssistBot Partly blockedNot explicitly addressed
GoogleOther Partly blockedNot explicitly addressed

How each site compares with its own rank peers

The sites may belong to different rank bands. Each percentage keeps its own population and denominator; Botfy does not subtract rates from unlike cohorts.

Traitdebate.com.mx hoteltonight.com
Wildcard group Yes 92.1% Ranks 10,001–100,000 · n=54,147 Yes 92.1% Ranks 10,001–100,000 · n=54,147
Effectively open file No 22.6% Ranks 10,001–100,000 · n=54,147 No 22.6% Ranks 10,001–100,000 · n=54,147
Sitemap declared No 61.7% Ranks 10,001–100,000 · n=54,147 Yes 61.7% Ranks 10,001–100,000 · n=54,147
robots.txt served Yes 59.7% Ranks 10,001–100,000 · n=91,507 Yes 59.7% Ranks 10,001–100,000 · n=91,507
robots.txt returned 404 No 7.9% Ranks 10,001–100,000 · n=91,507 No 7.9% Ranks 10,001–100,000 · n=91,507
Over the RFC parse limit No 0.2% Ranks 10,001–100,000 · n=54,147 No 0.2% Ranks 10,001–100,000 · n=54,147
Blanket disallow Yes 4.0% Ranks 10,001–100,000 · n=54,147 No 4.0% Ranks 10,001–100,000 · n=54,147
How this comparison is measured

Both sites use Botfy's latest published observations. A missing fetch, unresolved agent verdict, and measured absence are different states. Peer percentages use each site's exclusive Tranco rank band and print their own denominator. Populations, sources, and limits →

Keep comparing

Each suggestion replaces one side using an editorial collection or current rank proximity. Neither means the sites have similar policies.