cloudflare.com/llms.txt
Follows the format: a title, a one-line summary and organised sections, in a size a model can read at once.
Example gallery
Good, typical, broken and extreme llms.txt and robots.txt files from this month's Tranco top 1M. Every card says the rule that picked it, so nothing here is chosen by taste.
Follows the format: a title, a one-line summary and organised sections, in a size a model can read at once.
Follows the format: a title, a one-line summary and organised sections, in a size a model can read at once.
Closest to the median of 88 lines among files that follow the format.
Missing the title or the sections the format asks for.
Missing the title or the sections the format asks for.
The largest file that still follows the format.
The smallest file that still follows the format.
The highest-ranked site using the Shopify store template, the same layout as 2,203 other files.
The highest-ranked site using the WordPress SEO plugins, the same layout as 898 other files.
The highest-ranked site using the Wix AI-agent layout, the same layout as 74 other files.
The highest-ranked site using the Domain-for-sale page, the same layout as 42 other files.
Crawl-delay: 1e62Asks crawlers to wait more than 10⁴⁴ times the age of the universe between visits. Google ignores Crawl-delay entirely.
Disallow: … × 27,653A 1.7 MB file of individual Disallow lines. Google stops reading robots.txt after 500 KB.
User-agent: GPTBot
Disallow: /
… 40 moreTied for the most AI crawlers blocked by name in the top million.
User-agent: *
Disallow: /The highest-ranked sites in Chrome's real-visit data that tell every crawler to stay away: gstatic.com (#4), blogspot.com (#99), reddit.com (#109).
4,586,226 bytesThe largest complete robots.txt on the top million. Google reads only the first 500 KB. 75 more files stopped at a storage limit, so their full size is unknown.
User-agent: GPTBot
Disallow: /
User-agent: OAI-SearchBot
Allow: /The highest-ranked sites that block OpenAI's training crawler but let its search crawler in: one company, two decisions.
Picked by the rule on each card and screened so nothing unsuitable is showcased. Files can change after our crawl; links go to the live version. Real /llms.txt files found on sites in our top-500k sample; robots.txt files we could read, one per site on this month's Tranco top 1M.