BotForYou Updated 9 October 2026 · Common Crawl CC-MAIN-2026-39

Every AI bot · Scrapy

scrapy

A name associated with Scrapy whose purpose hasn't been clearly documented.

In robots.txt: User-agent: scrapy

The numbers

Named by1.6%of readable robots.txt files on the top 1M (9,396 of 600,979)
Of those, block it74.5%shut it out of the whole site (6,996 of 9,396)
Same sites, since CC-MAIN-2026-34+0.0 ptsshare of the same 394,690 sites that block it by name: 1.1% → 1.1%
Via Cloudflare's file2.0%of its blocks come from Cloudflare's ready-made robots.txt (141 of 6,996)
Block the whole site: 6,996Block some paths: 1,960Name it but allow it: 440
Every robots.txt file on the top 1M that names scrapy, by what the rules that apply to it say.

Who does what

The highest-ranked sites on the top million in each group. Click one to see its whole robots.txt verdict.

Block some paths

  1. zdf.de #3,090
  2. pressreader.com #5,094
  3. webtoons.com #5,202
  4. smartnews.com #8,237
  5. zdfheute.de #10,652

Name it and allow it

  1. termly.io #4,239
  2. techtimes.com #12,806
  3. taxfoundation.org #14,933
  4. zinio.com #16,277
  5. kling.ai #17,121

Lists skip sites we don't showcase (adult content).

What this means for you: its purpose isn't documented, so what blocking it changes is unclear. Decide each bot on what it's for, not who runs it.

Counts use one robots.txt per site on this month's Tranco top 1M (600,979 files we could read). Bot list: registry 2026-07-14.1, kept in sync with Dark Visitors and ai.robots.txt.