robots.txt & llms.txt across the Tranco top 1M
Data as of 2026-07-25 · robots.txt via Common Crawl CC-MAIN-2026-25 + our polite crawl · llms.txt via our polite crawl · Tranco list ID not recorded for this snapshot · crawls: 2026-09-07 · 2026-08-02 · 2026-07-25 · 2026-07-17 687,794 robots.txt fetches · 491,025 parsed · 54,000 llms.txt probes

Does your website platform set up robots.txt for you?

Some platforms ship a robots.txt you never see and cannot edit. Others leave it entirely to you. This is what each one actually does, measured across the Tranco top 1,000,000 on 2026-07-25 — not from platform documentation, but from their sites' real files.

Every platform label on this page is an inference, not a fact. We do not ask a site what it runs. We read the paths its robots.txt names — /wp-admin/ means WordPress, /commerce/digital-download/ means Squarespace — and score a confidence. That is a hypothesis with evidence behind it, and it is labelled hypothesis everywhere it appears.

Platform identification has not been published for this snapshot yet. Identifying a platform means re-reading every robots.txt in the corpus against a rule table — a separate pass from the crawl that produced the figures elsewhere on this site. It is running; this page fills in when it lands. No site is missing from the data because it lacks a platform, only because we have not looked yet.

How this was measured

Platform identification reads the allow and disallow paths in each site's robots.txt against a rule table, corroborated where possible by response headers — a Server: Squarespace header agreeing with Squarespace paths raises confidence; a header naming a different platform lowers it. Only high-confidence inferences are published. Full method on the methodology page.

Squarespace is a special case worth knowing: it serves one identical, uneditable robots.txt to every site it hosts, so its row describes a single file rather than an average over many. Platforms that leave the file to the site owner — WordPress among them — show real spread instead.

The robots.txt columns are measured on ranked domains with a parsed HTTP 200 robots.txt. The llms.txt column is measured on live members of the llms.txt probe cohort that also carry a platform inference — a smaller and differently selected population, as noted above.