robots.txt & llms.txt across the Tranco top 1M
Data as of 2026-08-02 · robots.txt via Common Crawl CC-MAIN-2026-30 + our polite crawl · llms.txt via our polite crawl · Tranco list L5684 (2026-07-14) · 1,000,000 panel domains · crawls: 2026-09-07 · 2026-08-02 · 2026-07-25 · 2026-07-17 1,000,000 robots.txt fetches · 608,834 parsed · 55,770 llms.txt probes

What does a good llms.txt look like?

Three real files from the crawl, side by side: what a typical file looks like, an exemplary one, and a pathological case. Selection is criteria-based, not editorial — the rules are stated on each card, so the data speaks. Bodies are captured snapshots (may differ from the live file today); we link out to the current version.

Typical

The median genuine file — what most publishers actually ship (near 88 lines).

No qualifying captured example this crawl.

Exemplary

Conformant and complete — a clear H1, a summary blockquote, well-organized sections, and a sane size. What 'good' looks like.

No qualifying captured example this crawl.

Pathological

A 200 that isn't a file — the anti-pattern our integrity filter excludes. Origins answer a missing path with a web page, their robots.txt, an error string, or a bare "OK", and every one of them returns 200.

No qualifying captured example this crawl.

Why this matters: publishing a path that looks present but returns something that isn't a file inflates naive adoption counts. Botfy counts a file only when the body actually contains a document — see how presence is decided.