BotForYou Updated 9 October 2026 · Common Crawl CC-MAIN-2026-39

Every AI bot · Common Crawl

CCBot

Common Crawl's crawler that collects web pages to train AI models.

In robots.txt: User-agent: CCBot

The numbers

Named by15.1%of readable robots.txt files on the top 1M (90,568 of 600,979)
Of those, block it60.7%shut it out of the whole site (55,010 of 90,568)
Same sites, since CC-MAIN-2026-34−0.6 ptsshare of the same 394,690 sites that block it by name: 9.7% → 9.1%
Via Cloudflare's file49.0%of its blocks come from Cloudflare's ready-made robots.txt (26,974 of 55,010)Cloudflare stopped adding that file on 16 September 2026, so this month mixes before and after. What happened
Block the whole site: 55,010Block some paths: 6,617Name it but allow it: 28,941
Every robots.txt file on the top 1M that names CCBot, by what the rules that apply to it say.

Who does what

The highest-ranked sites on the top million in each group. Click one to see its whole robots.txt verdict.

Block some paths

  1. reg.ru #142
  2. android.com #256
  3. linktr.ee #272
  4. amplitude.com #298
  5. indeed.com #371

Name it and allow it

  1. cloudflare.com #2
  2. wordpress.org #48
  3. ui.com #110
  4. kaspersky.com #140
  5. cisco.com #253

Lists skip sites we don't showcase (adult content).

What this means for you: blocking CCBot asks Common Crawl not to collect your pages for AI training from now on. robots.txt only affects future visits, and it doesn't touch Common Crawl's other bots. Decide each bot on what it's for, not who runs it.

Counts use one robots.txt per site on this month's Tranco top 1M (600,979 files we could read). Bot list: registry 2026-07-14.1, kept in sync with Dark Visitors and ai.robots.txt.