AI crawlers in 2026: who's who, what robots.txt can actually block, and what it can't

"Block AI bots in robots.txt" is advice you'll find almost everywhere. What most sources leave out: half the bots on those lists don't have a crawler at all, and another chunk simply ignores robots.txt. This overview is built on each company's own primary documentation (see Sources), not secondary summaries, and it's the same list our audit tool's ACC-01 check uses (config/ai-bots.json).

Who's who

TokenOperatorPurposeRespects robots.txt?
GPTBotOpenAIModel trainingYes
OAI-SearchBotOpenAIIndexing for ChatGPT searchYes
ChatGPT-UserOpenAIReal-time fetch on user request"May not apply" (user-initiated)
ClaudeBotAnthropicModel trainingYes
Claude-User / Claude-SearchBotAnthropicReal-time fetch / search indexingYes
PerplexityBotPerplexityIndexingYes
Perplexity-UserPerplexityReal-time fetch on user requestNo, documented
Google-ExtendedGoogleTraining opt-out tokenHas no crawler of its own
Applebot-ExtendedAppleTraining opt-out tokenHas no crawler of its own
CCBotCommon CrawlTraining datasetYes, including non-standard Crawl-delay
BytespiderByteDanceTrainingUnreliably
Meta-ExternalAgentMetaTraining and indexingYes
Meta-ExternalFetcherMetaUser-initiated fetch (agentic AI)May bypass
DuckAssistBotDuckDuckGoFetch for DuckAssistYes, opt-out takes up to 72h to apply

The full list, including the three Mistral tokens (MistralAI-User, MistralAI-Index, MistralAI-Training) and Amazonbot, lives in our app's config/ai-bots.json, in the exact format the audit evaluates.

Two tokens that aren't crawlers

Google-Extended and Applebot-Extended show up in lists alongside real crawlers, but they're a different category. Google states it plainly: "Google-Extended doesn't have a separate HTTP request user agent string. Crawling is done with existing Google user agent strings." It's purely a robots.txt token that controls whether data already collected by regular Googlebot may be used to train Gemini. Google adds that Google-Extended does not impact a site's inclusion in Google Search nor is it used as a ranking signal.

Applebot-Extended works the same way, because it is a secondary token on top of the regular Applebot that only governs whether already-collected data may be used to train Apple Intelligence. Blocking it doesn't affect Siri or Spotlight. Describing either token as a "bot that visits your site" is inaccurate. Googlebot or Applebot does the visiting; the token only changes what the data may be used for afterward.

What robots.txt can't reliably enforce

For three crawlers, a robots.txt entry is a weak or non-functional control:

  • Perplexity-User documentedly ignores robots.txt, and Perplexity says so in its own documentation. Cloudflare also reported undeclared Perplexity crawlers with rotating user agents and IPs in August 2025.
  • Meta-ExternalFetcher is a user-initiated fetch (unlike Meta-ExternalAgent), and Meta itself documents that it may bypass robots.txt.
  • Bytespider (ByteDance) has no primary source at all, because the company publishes neither documentation nor a verification mechanism (IP/DNS). Its user agent string is known only from observed traffic, so a UA match is a clue, not proof of identity. Independent monitoring has repeatedly recorded crawling on URLs disallowed in robots.txt.

For Bytespider, a robots.txt entry alone isn't enough, because reliable enforcement needs server-side blocking (rate limiting, a WAF, or IP-range blocking), not just a file the bot apparently doesn't read anyway.

What Anthropic does (and doesn't) publish

Anthropic doesn't officially publish user agent strings for ClaudeBot, Claude-User or Claude-SearchBot. The file at claude.com/crawling/bots.json is a shared IP allowlist, so it confirms a request came from Anthropic, but doesn't distinguish which specific bot sent it. The legacy tokens Claude-Web and anthropic-ai don't appear in the current support article, so don't rely on them.

What an entry actually looks like

The syntax is the same as any other robots.txt entry: a User-agent block with a token from the table above and Disallow/Allow underneath. A typical pattern blocks training while letting through a real-time fetch made on a specific user's behalf:

  • User-agent: GPTBot / Disallow: / blocks OpenAI's model training across the whole site.
  • User-agent: ChatGPT-User / Allow: / lets through the fetch ChatGPT makes at a specific user's explicit request (typically when someone asks about a particular page).
  • User-agent: Bytespider / Disallow: /: add this entry anyway, even knowing it isn't reliable on its own; it's a first layer, not the last one.

A block for Google-Extended or Applebot-Extended has the same shape, but it doesn't stop any visit at all, because it only says data already collected by regular Googlebot/Applebot may not be used for training. To stop the crawling itself you'd have to block Googlebot/Applebot directly, which also affects classic search.

How to verify a request is really from that company

A user agent string is self-declared, so anyone can set one. OpenAI, Anthropic and Google all publish IP ranges or rDNS records you can check traffic against (Anthropic's is claude.com/crawling/bots.json, mentioned above). Before writing a robots.txt rule based purely on a user agent name seen in your logs, verify at least a sample of requests against the published IP range. For Bytespider this isn't possible at all, since ByteDance publishes no such list.

What to do about it

Deciding which bots to allow and which to block is a business decision (training, indexing, and on-demand fetching are three different things), not a technical necessity. Typical scenarios: an online store that doesn't want its pricing training a competitor's model blocks the training bots (GPTBot, ClaudeBot, Meta-ExternalAgent) but lets indexing and fetch bots through, because it wants to show up in AI answers. A publisher whose business depends on ad revenue against its own content might block indexing too. A documentation site typically welcomes every bot, because being cited as often as possible is the whole point.

What's worth doing regardless of the strategy you choose: keeping robots.txt current and readable, not expecting any traffic effect from opt-out tokens, and treating unreliable bots (Bytespider) as requiring more than a robots.txt line.

Torumata's audit (check ACC-01) compares your robots.txt against the current list and generates a patch. For each bot you'll see whether the entry is actually enforceable or purely declarative. What crawlers can even see in the first place is covered in JavaScript and AI crawlers.

Want to see where your own robots.txt stands today? Run a free audit.

Sources