robots.txt & llms.txt across the Tranco top 1M
Data as of 2026-08-02 · robots.txt via Common Crawl CC-MAIN-2026-30 + our polite crawl · llms.txt via our polite crawl · Tranco list L5684 (2026-07-14) · 1,000,000 panel domains · crawls: 2026-09-07 · 2026-08-02 · 2026-07-25 · 2026-07-17 1,000,000 robots.txt fetches · 608,834 parsed · 55,770 llms.txt probes

Does your website platform set up robots.txt for you?

Some platforms ship a robots.txt you never see and cannot edit. Others leave it entirely to you. This is what each one actually does, measured across the Tranco top 1,000,000 on 2026-08-02 — not from platform documentation, but from their sites' real files.

Every platform label on this page is an inference, not a fact. We do not ask a site what it runs. We read the paths its robots.txt names — /wp-admin/ means WordPress, /commerce/digital-download/ means Squarespace — and score a confidence. That is a hypothesis with evidence behind it, and it is labelled hypothesis everywhere it appears.

What each platform's robots.txt does

Platform Sites measured Block every crawler Allow every crawler Declare a sitemap
WordPress 61,301 0.3% 1.3% 84.2%
Shopify 29,105 0.1% 0.4% 98.5%
Drupal 11,467 0.4% 1.5% 24.8%
Cloudflare 7,781 0.2% 0.9% 82.0%
Joomla 5,413 0.8% 1.4% 27.2%
Wix 1,278 0.1% 2.0% 99.4%
Next.js 881 1.0% 2.0% 95.3%
Magento 485 0.6% 3.3% 85.8%
Squarespace 52 0.0% 1.9% 98.1%
Ghost 11 0.0% 0.0% 81.8%

"Sites measured" counts ranked domains whose robots.txt carried enough signal to identify the platform. A site can match more than one row — a WordPress site behind Cloudflare appears in both, which is correct: they are different layers, not competing answers.

Does your platform publish an llms.txt for you?

Large differences, and not the ones you would guess from platform marketing.

Platform Live sites in the probe cohort Publish an llms.txt
WordPress 3,702 19.3%
Shopify 1,328 93.8%
Drupal 1,281 6.6%
Cloudflare 940 20.6%
Joomla 383 6.3%
Wix 69 too few to report
Next.js 95 too few to report
Magento 25 too few to report
Squarespace 9 too few to report
Ghost 4 too few to report

A rate is published only for platforms with at least 100 live sites in the probe cohort. Below that a percentage is noise wearing a percentage sign — so the count is shown and the rate withheld, rather than printing a figure that would move by tens of points on one more site.

The high rates are platforms generating the file, not owners writing it. Where a platform ships llms.txt as a default, nearly every store on it has one — and the files are identical apart from the name. Across measured Shopify stores that publish an llms.txt, 93% open with the same generated heading (Agent Instructions — <store name>). That is one decision made once by a platform, showing up as thousands of sites "adopting" a convention.

It matters beyond this page: any llms.txt adoption figure — ours included — counts those files. Adoption is not the same as anyone choosing to adopt.

Do not compare these to the site-wide llms.txt figure. They are measured on a different, self-selected population and they run higher for a reason that has nothing to do with the platforms. A site only earns a platform label here if its robots.txt names distinctive paths — which means somebody configured it. Configured sites adopt llms.txt more than unconfigured ones. These rates compare platforms to each other, which is the question this page answers; they are not a read on the web at large.

How this was measured

Platform identification reads the allow and disallow paths in each site's robots.txt against a rule table, corroborated where possible by response headers — a Server: Squarespace header agreeing with Squarespace paths raises confidence; a header naming a different platform lowers it. Only high-confidence inferences are published. Full method on the methodology page.

Squarespace is a special case worth knowing: it serves one identical, uneditable robots.txt to every site it hosts, so its row describes a single file rather than an average over many. Platforms that leave the file to the site owner — WordPress among them — show real spread instead.

The robots.txt columns are measured on ranked domains with a parsed HTTP 200 robots.txt. The llms.txt column is measured on live members of the llms.txt probe cohort that also carry a platform inference — a smaller and differently selected population, as noted above.