Every AI bot · img2dataset
img2dataset
img2dataset's crawler that collects web pages to train AI models.
In robots.txt: User-agent: img2dataset
The numbers
57.4%
41.2%
1.4%
Block the whole site: 4,229Block some paths: 3,038Name it but allow it: 102
Who does what
The highest-ranked sites on the top million in each group. Click one to see its whole robots.txt verdict.
Block it
- amazon.com #24
- yahoo.com #59
- msn.com #69
- chatgpt.com #75
- amazonvideo.com #123
Block some paths
- indeed.com #371
- squarespace.com #544
- deezer.com #1,828
- web.de #2,479
- gmx.net #2,498
Name it and allow it
- lachainemeteo.com #8,132
- zinio.com #16,277
- kling.ai #17,121
- biu.ac.il #29,152
- meteoconsult.fr #33,558
Lists skip sites we don't showcase (adult content).
What this means for you: blocking img2dataset asks img2dataset not to collect your pages for AI training from now on. robots.txt only affects future visits, and it doesn't touch img2dataset's other bots. Decide each bot on what it's for, not who runs it.
Counts use one robots.txt per site on this month's Tranco top 1M (600,979 files we could read). Bot list: registry 2026-07-14.1, kept in sync with Dark Visitors and ai.robots.txt.