BotForYou Updated 9 October 2026 · Common Crawl CC-MAIN-2026-39

Every AI bot · Apple

Applebot-Extended

Not a crawler. It's a name Apple reads in robots.txt to decide whether pages it collects for other reasons may be used to train AI. Blocking it doesn't stop Apple's other crawlers.

In robots.txt: User-agent: Applebot-Extended

The numbers

Named by8.8%of readable robots.txt files on the top 1M (52,680 of 600,979)
Of those, block it80.3%shut it out of the whole site (42,312 of 52,680)
Same sites, since CC-MAIN-2026-34−0.6 ptsshare of the same 394,690 sites that block it by name: 7.8% → 7.2%
Via Cloudflare's file63.7%of its blocks come from Cloudflare's ready-made robots.txt (26,961 of 42,312)Cloudflare stopped adding that file on 16 September 2026, so this month mixes before and after. What happened
Block the whole site: 42,312Block some paths: 6,344Name it but allow it: 4,024
Every robots.txt file on the top 1M that names Applebot-Extended, by what the rules that apply to it say.

Apple's other bots

Sites often treat one company's bots differently, depending on what each is for.

BotPurposeFiles naming itOf those, block it
ApplebotBuilds an AI search index37,76711.3%

Who does what

The highest-ranked sites on the top million in each group. Click one to see its whole robots.txt verdict.

Block some paths

  1. facebook.com #3
  2. oracle.com #205
  3. ubuntu.com #233
  4. android.com #256
  5. amplitude.com #298

Name it and allow it

  1. wordpress.org #48
  2. ui.com #110
  3. kaspersky.com #140
  4. forter.com #270
  5. linktr.ee #272

Lists skip sites we don't showcase (adult content).

What this means for you: blocking Applebot-Extended is how you tell Apple not to use your pages for AI training, without leaving Apple's search. Decide each bot on what it's for, not who runs it.

Counts use one robots.txt per site on this month's Tranco top 1M (600,979 files we could read). Bot list: registry 2026-07-14.1, kept in sync with Dark Visitors and ai.robots.txt.