BotForYou Updated 9 October 2026 · Common Crawl CC-MAIN-2026-39

Every AI bot · Allen Institute

ai2bot-dolma

Allen Institute's crawler that collects web pages to train AI models.

In robots.txt: User-agent: ai2bot-dolma

The numbers

Named by1.2%of readable robots.txt files on the top 1M (7,022 of 600,979)
Of those, block it55.0%shut it out of the whole site (3,862 of 7,022)
Same sites, since CC-MAIN-2026-34+0.0 ptsshare of the same 394,690 sites that block it by name: 0.6% → 0.7%
Via Cloudflare's file1.5%of its blocks come from Cloudflare's ready-made robots.txt (59 of 3,862)
Block the whole site: 3,862Block some paths: 3,042Name it but allow it: 118
Every robots.txt file on the top 1M that names ai2bot-dolma, by what the rules that apply to it say.

Allen Institute's other bots

Sites often treat one company's bots differently, depending on what each is for.

BotPurposeFiles naming itOf those, block it
ai2botCollects pages to train AI10,13561.2%

Who does what

The highest-ranked sites on the top million in each group. Click one to see its whole robots.txt verdict.

Block some paths

  1. squarespace.com #544
  2. deezer.com #1,828
  3. pressreader.com #5,094
  4. youradchoices.ca #6,541
  5. qwant.com #7,038

Name it and allow it

  1. lachainemeteo.com #8,132
  2. observador.pt #14,136
  3. downdetector.com #15,673
  4. zinio.com #16,277
  5. rojgarresult.com #23,235

Lists skip sites we don't showcase (adult content).

What this means for you: blocking ai2bot-dolma asks Allen Institute not to collect your pages for AI training from now on. robots.txt only affects future visits, and it doesn't touch Allen Institute's other bots. Decide each bot on what it's for, not who runs it.

Counts use one robots.txt per site on this month's Tranco top 1M (600,979 files we could read). Bot list: registry 2026-07-14.1, kept in sync with Dark Visitors and ai.robots.txt.