Every AI bot · Scrapy
scrapy
A name associated with Scrapy whose purpose hasn't been clearly documented.
In robots.txt: User-agent: scrapy
The numbers
74.5%
20.9%
4.7%
Block the whole site: 6,996Block some paths: 1,960Name it but allow it: 440
Who does what
The highest-ranked sites on the top million in each group. Click one to see its whole robots.txt verdict.
Block it
- facebook.com #3
- instagram.com #11
- fbcdn.net #13
- linkedin.com #17
- amazon.com #24
Block some paths
- zdf.de #3,090
- pressreader.com #5,094
- webtoons.com #5,202
- smartnews.com #8,237
- zdfheute.de #10,652
Name it and allow it
- termly.io #4,239
- techtimes.com #12,806
- taxfoundation.org #14,933
- zinio.com #16,277
- kling.ai #17,121
Lists skip sites we don't showcase (adult content).
What this means for you: its purpose isn't documented, so what blocking it changes is unclear. Decide each bot on what it's for, not who runs it.
Counts use one robots.txt per site on this month's Tranco top 1M (600,979 files we could read). Bot list: registry 2026-07-14.1, kept in sync with Dark Visitors and ai.robots.txt.