ACC-01
robots.txt blocks AI crawlers
What the check measures
The audit fetches /robots.txt, splits it into groups by User-agent, and compares those against the list of AI bots we track. The list lives in configuration, not in code: config/ holds 25 tokens from OpenAI, Anthropic, Google, Perplexity, Apple, Amazon, Meta, Mistral, DuckDuckGo, Common Crawl, ByteDance and Exa, each with its operator, its purpose (training / indexing / user-initiated fetch) and whether writing it into robots.txt even means anything.
The finding is raised only when a group blocks the whole site: it contains Disallow: / or Disallow: /* and there is no Allow: / in that same group. That precedence is not our convention but the standard's rule (RFC 9309: the longest matching rule wins, and on a tie Allow wins).
What the check does not do:
- Partial blocks are never reported.
Disallow: /blog/orDisallow: /*.pdfraises nothing, even though it may shut off a large part of your content. This check judges only a whole-site ban. - It does not verify that a bot obeys the rule. It reads your file, not another company's behaviour. For three tokens it is in fact documented that they need not obey; see below.
- It never looks at
<meta name="robots">or theX-Robots-Tagheader. That isACC-07, and it can drop a page out of the index with a flawlessrobots.txt. - It cannot see blocking at the CDN or firewall level, which leaves no trace in
robots.txtat all. That isACC-09. - Bots absent from our configuration are ignored. A ban on an unknown agent shows up nowhere.
- It does not work out inherited rules. We report only bots your file names. One of them, however, also follows another group's rules; see the paragraph below.
One bot inherits search-engine rules: ExaSearchBot. Exa Labs runs an index that developer tools and agents draw on (Cursor, Devin and others). Its own documentation (verified 17 Sept 2026 at crawler.exa.ai) states that if you have not given it a group of its own, it follows whatever rules you set for the major search engines, and only then the * group. In practice that means User-agent: Googlebot + Disallow: / shuts Exa out as well, without your file ever mentioning it. This check stays silent about it, because it reports only tokens named explicitly. Such a site does get the critical ACC-10 finding, and fixing robots.txt lifts both bans at once. If you want certainty, give Exa a group of its own.
The finding is site-level, appears at most once in the report, and lists the specific tokens we found in your file, each with a note on what blocking it actually means.
How strong the evidence is
Without this, AI search has no way to show you at all. Google states it in its own documentation: a page must be indexed and eligible to be shown with a snippet. The strongest class we have.
This is the firmest thing the entire audit can tell you. It rests on neither measurement nor estimate, but on a sentence from Google's own documentation:
To be eligible to be shown as a supporting link in AI Overviews or AI Mode, a page must be indexed and eligible to be shown in Google Search with a snippet. There are no additional requirements.Google Search Central, AI features
If a crawler is not allowed in, it does not index you, and everything else in the report becomes moot. The same holds for Microsoft Copilot, which builds on Bing's index. That is why ACC-01 is a critical finding rather than a suggestion: it is not an optimisation, it is the ticket in.
Where we have to be precise, and where we part ways with guides that point at robots.txt indiscriminately: for four tokens, writing a rule does not mean what people expect.
| Token | What is different | What follows |
|---|---|---|
Google-Extended | It has no crawler and no user agent of its own. Google says literally that crawling is done with existing Google user agent strings. | It governs only whether data is used to train Gemini. It affects neither inclusion nor ranking in Search; Google states that explicitly. |
Applebot-Extended | It does not crawl either; it is a secondary token layered over Applebot. | Blocking it affects neither Siri nor Spotlight, only the use of data for Apple Intelligence. |
Perplexity-User | Documented as not respecting robots.txt (a fetch triggered by a user's question). | A rule in the file will not fix it. Promising otherwise would be untrue. |
Bytespider (ByteDance) | ByteDance publishes neither documentation nor an IP/DNS verification mechanism. | The rule is worth writing, but real enforcement needs server-side blocking (rate limiting or a WAF). |
A fifth caveat belongs to Meta-: unlike Meta- it is a user-initiated fetch, and those may bypass robots.txt. And DuckAssistBot behaves correctly, just slowly: its opt-out takes about 72 hours to take effect, so nothing being different the next day is not a fault.
One last thing worth saying out loud: blocking training crawlers can be a legitimate business decision, not a mistake. A publisher who does not want their archive used for free to train a model is doing something entirely reasonable. So the finding does not say “you have a bug”; it says what it costs you and leaves the decision to you. The critical severity is there because in the overwhelming majority of cases we see, the block was never intended: inherited from a template, left over from a staging deployment, or a blanket Disallow nobody remembers writing.
How to fix it
Open robots.txt at the root of your site and find the group that bans the whole site. The usual culprit is a single block left over from the days before launch:
User-agent: *
Disallow: /If you do want to let crawlers in, add explicit rules for the ones you want. Each group stands alone and inherits nothing from the others:
User-agent: GPTBot
Allow: /
User-agent: OAI-SearchBot
Allow: /
User-agent: ClaudeBot
Allow: /
User-agent: PerplexityBot
Allow: /
Sitemap: https://your-domain/sitemap.xmlA few places where this goes wrong in practice:
- One group, one bot.
User-agent: GPTBot, ClaudeBoton a single line is not part of the standard; write two groups. - A blank line separates groups. A rule written below a blank line belongs to a different group than the one you meant it for.
- A specific group overrides
User-agent: *. If a bot has its own group, it never reads the wildcard one; addingAllowto the wildcard group is not enough when the bot also has a group of its own. - The file must sit at the domain root (
https://) and return 200. A subdomain is governed by its own file, not the apex one.your- domain/ robots. txt - After editing, open the file in a browser. Plugins and frameworks often generate it, so what is in your repository need not be what is served.
If you want to keep crawlers out and know why, leave in only the tokens you genuinely mean to block, and account for the four exceptions in the table above. For Bytespider, add server-side blocking as well; the rule alone has no way to take effect.
What the report says about it
Finding description
robots.txt contains a Disallow rule for these monitored AI user-agents: [the list of blocked bots]. If these are actual crawlers (not opt-out tokens without their own UA/IP, such as Google-
Recommendation
Add explicit rules to robots.txt for the AI crawlers you want to allow, for example:
User-agent: GPTBot
Allow: /
User-agent: ClaudeBot
Allow: /
Keep Disallow only for user-agents you deliberately want to block. Watch for a few exceptions where a robots.txt entry alone doesn't do what you'd expect: Google-
Sources
- Google Search Central: AI features and your website (accessed 2026-08-19)
- Google: Common crawlers (Google-Extended) (accessed 2026-08-02)
- OpenAI: Overview of OpenAI crawlers (accessed 2026-08-19)
- Anthropic Support: Does Anthropic crawl data from the web? (accessed 2026-08-19)
- Perplexity: Crawlers (accessed 2026-08-02)
- Apple Support: About Applebot (accessed 2026-08-02)
Text verified 2026-09-12