What does a good llms.txt look like?
Three real files from the crawl, side by side: what a typical file looks like, an exemplary one, and a pathological case. Selection is criteria-based, not editorial — the rules are stated on each card, so the data speaks. Bodies are captured snapshots (may differ from the live file today); we link out to the current version.
The median genuine file — what most publishers actually ship (near 88 lines).
No qualifying captured example this crawl.
Conformant and complete — a clear H1, a summary blockquote, well-organized sections, and a sane size. What 'good' looks like.
No qualifying captured example this crawl.
A 200 that isn't a file — the anti-pattern our integrity filter excludes. Origins answer a missing path with a web page, their robots.txt, an error string, or a bare "OK", and every one of them returns 200.
No qualifying captured example this crawl.
Why this matters: publishing a path that looks present but returns something that isn't a file inflates naive adoption counts. Botfy counts a file only when the body actually contains a document — see how presence is decided.