Every AI bot · Anthropic
ClaudeBot
Anthropic's crawler that collects web pages to train AI models.
In robots.txt: User-agent: ClaudeBot
The numbers
Anthropic's other bots
Sites often treat one company's bots differently, depending on what each is for.
| Bot | Purpose | Files naming it | Of those, block it |
|---|---|---|---|
| anthropic-ai | Purpose not documented | 28,018 | 62.2% |
| claude-web | Fetches a page for a chatbot user | 21,076 | 65.7% |
| Claude-SearchBot | Builds an AI search index | 10,987 | 36.2% |
| Claude-User | Fetches a page for a chatbot user | 9,951 | 32.9% |
Who does what
The highest-ranked sites on the top million in each group. Click one to see its whole robots.txt verdict.
Block it
- instagram.com #11
- fbcdn.net #13
- linkedin.com #17
- amazon.com #24
- tiktok.com #54
Block some paths
- facebook.com #3
- github.com #29
- netflix.com #41
- oracle.com #205
- ubuntu.com #233
Name it and allow it
- digicert.com #47
- wordpress.org #48
- ui.com #110
- kaspersky.com #140
- godaddy.com #154
Lists skip sites we don't showcase (adult content).
What this means for you: blocking ClaudeBot asks Anthropic not to collect your pages for AI training from now on. robots.txt only affects future visits, and it doesn't touch Anthropic's other bots. Decide each bot on what it's for, not who runs it.
Counts use one robots.txt per site on this month's Tranco top 1M (600,979 files we could read). Bot list: registry 2026-07-14.1, kept in sync with Dark Visitors and ai.robots.txt.