AI crawlers in 2026: who's who, what robots.txt can actually block, and what it can't
"Block AI bots in robots.txt" is advice you'll find almost everywhere. What most sources leave out: half the bots on those lists don't have a crawler at all, and another chunk simply ignores robots.txt. This overview is built on each company's own primary documentation (see Sources), not secondary summaries, and it's the same list our audit tool's ACC-01 check uses (config/).
Who's who
| Token | Operator | Purpose | Respects robots.txt? |
|---|---|---|---|
GPTBot | OpenAI | Model training | Yes |
OAI-SearchBot | OpenAI | Indexing for ChatGPT search | Yes |
ChatGPT-User | OpenAI | Real-time fetch on user request | "May not apply" (user- |
ClaudeBot | Anthropic | Model training | Yes |
Claude-User / Claude-SearchBot | Anthropic | Real-time fetch / search indexing | Yes |
PerplexityBot | Perplexity | Indexing | Yes |
Perplexity-User | Perplexity | Real-time fetch on user request | No, documented |
Google-Extended | Training opt-out token | Has no crawler of its own | |
Applebot-Extended | Apple | Training opt-out token | Has no crawler of its own |
CCBot | Common Crawl | Training dataset | Yes, including non-standard Crawl-delay |
Bytespider | ByteDance | Training | Unreliably |
Meta-ExternalAgent | Meta | Training and indexing | Yes |
Meta-ExternalFetcher | Meta | User-initiated fetch (agentic AI) | May bypass |
DuckAssistBot | DuckDuckGo | Fetch for DuckAssist | Yes, opt-out takes up to 72h to apply |
The full list, including the three Mistral tokens (MistralAI-User, MistralAI-, MistralAI-) and Amazonbot, lives in our app's config/, in the exact format the audit evaluates.
Two tokens that aren't crawlers
Google- and Applebot- show up in lists alongside real crawlers, but they're a different category. Google states it plainly: "Google-Google- does not impact a site's inclusion in Google Search nor is it used as a ranking signal.
Applebot- works the same way, because it is a secondary token on top of the regular Applebot that only governs whether already-
What robots.txt can't reliably enforce
For three crawlers, a robots.txt entry is a weak or non-functional control:
Perplexity-documentedly ignoresUser robots.txt, and Perplexity says so in its own documentation. Cloudflare also reported undeclared Perplexity crawlers with rotating user agents and IPs in August 2025.Meta-is a user-initiated fetch (unlikeExternalFetcher Meta-), and Meta itself documents that it may bypassExternalAgent robots.txt.Bytespider(ByteDance) has no primary source at all, because the company publishes neither documentation nor a verification mechanism (IP/DNS). Its user agent string is known only from observed traffic, so a UA match is a clue, not proof of identity. Independent monitoring has repeatedly recorded crawling on URLs disallowed inrobots.txt.
For Bytespider, a robots.txt entry alone isn't enough, because reliable enforcement needs server-side blocking (rate limiting, a WAF, or IP-range blocking), not just a file the bot apparently doesn't read anyway.
What Anthropic does (and doesn't) publish
Anthropic doesn't officially publish user agent strings for ClaudeBot, Claude-User or Claude-. The file at claude. is a shared IP allowlist, so it confirms a request came from Anthropic, but doesn't distinguish which specific bot sent it. The legacy tokens Claude-Web and anthropic-ai don't appear in the current support article, so don't rely on them.
What an entry actually looks like
The syntax is the same as any other robots.txt entry: a User-agent block with a token from the table above and Disallow/Allow underneath. A typical pattern blocks training while letting through a real-time fetch made on a specific user's behalf:
User-agent: GPTBot/Disallow: /blocks OpenAI's model training across the whole site.User-agent: ChatGPT-User/Allow: /lets through the fetch ChatGPT makes at a specific user's explicit request (typically when someone asks about a particular page).User-agent: Bytespider/Disallow: /: add this entry anyway, even knowing it isn't reliable on its own; it's a first layer, not the last one.
A block for Google- or Applebot- has the same shape, but it doesn't stop any visit at all, because it only says data already collected by regular Googlebot/Googlebot/Applebot directly, which also affects classic search.
How to verify a request is really from that company
A user agent string is self-declared, so anyone can set one. OpenAI, Anthropic and Google all publish IP ranges or rDNS records you can check traffic against (Anthropic's is claude., mentioned above). Before writing a robots.txt rule based purely on a user agent name seen in your logs, verify at least a sample of requests against the published IP range. For Bytespider this isn't possible at all, since ByteDance publishes no such list.
What to do about it
Deciding which bots to allow and which to block is a business decision (training, indexing, and on-demand fetching are three different things), not a technical necessity. Typical scenarios: an online store that doesn't want its pricing training a competitor's model blocks the training bots (GPTBot, ClaudeBot, Meta-) but lets indexing and fetch bots through, because it wants to show up in AI answers. A publisher whose business depends on ad revenue against its own content might block indexing too. A documentation site typically welcomes every bot, because being cited as often as possible is the whole point.
What's worth doing regardless of the strategy you choose: keeping robots.txt current and readable, not expecting any traffic effect from opt-out tokens, and treating unreliable bots (Bytespider) as requiring more than a robots.txt line.
Torumata's audit (check ACC-01) compares your robots.txt against the current list and generates a patch. For each bot you'll see whether the entry is actually enforceable or purely declarative. What crawlers can even see in the first place is covered in JavaScript and AI crawlers.
Want to see where your own robots.txt stands today? Run a free audit.
Sources
- OpenAI: Overview of OpenAI crawlers (accessed 2026-08-02)
- Anthropic Support: Does Anthropic crawl data from the web? (accessed 2026-08-02)
- Perplexity: Crawlers documentation (accessed 2026-08-02)
- Google: common crawlers and fetchers (accessed 2026-08-02)
- Common Crawl: CCBot (accessed 2026-08-02)
- Meta: web crawlers (accessed 2026-08-02)
- Mistral: robots (accessed 2026-08-02)
- Apple: Applebot-Extended (accessed 2026-08-02)