ACC-10

robots.txt blocks Googlebot from the entire site

CriticalCrawler access

What the check measures

In robots.txt (the same file the audit fetches for ACC-01) we look for one thing only: whether it denies Googlebot the entire site. That is, a Disallow: / rule in the group that applies to Googlebot.

User-agent: Googlebot
Disallow: /

The group is selected by user-agent specificity, exactly as RFC 9309 describes it and as a real crawler would resolve it. That is the crux of this check, so here it is in order:

  • Does Googlebot have its own group in the file? Then that group alone decides, and the User-agent: * group is ignored for it entirely, even if that group denies everything. So a file with Disallow: / under * and a User-agent: Googlebot group with Allow: / below it is NOT a blocked site, and no finding comes out of it.
  • No group of its own? Then User-agent: * applies to it. If that group denies the root, Googlebot is blocked just as effectively as if it had been named. In that case the report says the * rule is what decided.
  • The token is matched in full, not by prefix. Googlebot-Image is a different crawler. A site that only disallowed image indexing will not get this finding.
  • Allow: / in the same group cancels the root disallow. When rules are of equal length, Allow wins (RFC 9309).

What the check does not do:

  • It does not assess partial path disallows. Disallow: /admin/ for Googlebot is normal and produces no finding. This is solely about denying the WHOLE site.
  • It does not verify that you are actually out of the index. That would mean asking Google about your site from the outside; we only read your own file. The finding reports what the file says, not what Google holds in its database.
  • It does not cover other search engines. We do not track Bingbot or Applebot; the next section explains why.
  • It has nothing to do with AI crawlers. Those are ACC-01, and Googlebot is not one of them. That these are two different things is the entire reason this check has its own code.

The finding is site-wide and its severity is critical. The list of tracked tokens is not in the code; it lives in config/ai-bots.json under search_gatekeepers, kept separate from the list of AI crawlers.

How strong the evidence is

Necessary condition

Without this, AI search has no way to show you at all. Google states it in its own documentation: a page must be indexed and eligible to be shown with a snippet. The strongest class we have.

This is the strongest claim in the whole audit, and the only place where we talk about a condition rather than a probability. Google states it in its own documentation:

“To be eligible to be shown as a supporting link in AI Overviews or AI Mode, a page must be indexed and eligible to be shown in Google Search with a snippet. There are no additional requirements.”

Read that carefully, because the whole finding rests on it. It is not only about a page being indexed, but about a page being eligible to be shown with a snippet. And eligibility is exactly what Disallow: / takes away.

A common objection: “but a disallowed page can still end up in the index”. That is true, and it changes nothing. When links from elsewhere point at a disallowed URL, Google may record it in the index, but only as a bare URL with no content, because it was not allowed to fetch it. So no snippet is shown for it, and the second half of the condition is not met.

Why only Googlebot and not Bingbot. Microsoft Copilot builds on the Bing index, so the same ought to hold, and docs/externi-zavislosti.md §4c does infer exactly that from the architecture. But an inference from architecture is not the same as a sentence in the documentation, and a claim of this class may rest only on a primary source. We have no statement from Microsoft to that effect, so we do not track Bingbot. Once such a statement is found, a token gets added to the config; no new check needs to be written for it.

And a note that belongs here even though it does not sell well: denying Googlebot is quite often the right decision. Staging and client preview sites, intranets, any environment that is not meant to reach search: that is how it should be there. So the finding is not a reproach but an observation. If you know why it is there, “yes, that is what I want” is a perfectly valid answer, and the report ends there for you.

How to fix it

First establish whether it is intentional. Open https://your-domain/robots.txt in a browser and look at what is actually there, not what is in your repository or what you remember from setup. The most common cause of an unexpected block is a site that was built on a staging address and whose robots.txt was never switched back at launch.

If it is intentional, change nothing. Just be aware that such a site will not appear in Google Search or in its AI answers, and that no other finding in this audit changes that. While this block stands, work on content and structure is an investment in something nobody will see.

If it is not intentional, you have two routes:

SituationWhat to do
Googlebot has its own group with Disallow: /Delete that line from it, or replace it with Allow: /.
The block comes from the User-agent: * groupEither remove Disallow: / there, or give Googlebot its own group with Allow: /.

Things to watch out for:

  • A group naming Googlebot overrides User-agent: * AS A WHOLE, not just on that one line. This is the most common mistake when making this fix: you give Googlebot a group with Allow: /, and in doing so you also cancel every partial disallow that used to apply to it through *. If Disallow: /admin/ should still apply, repeat it inside the new group.
  • Do not confuse this with the meta tag. robots.txt says where a crawler may go; <meta name="robots" content="noindex"> says what it may index. They are two different places handled by two different checks (the latter is ACC-07). A block in robots.txt also makes a page's noindex impossible to read: the crawler never reaches it.
  • Check your CDN too. Some hosts and CDNs assemble robots.txt themselves, and the file in your repository may not be the one served to visitors. ACC-09 covers the risk from a CDN's default settings.

After deploying, verify the change in Google Search Console: Settings shows the fetched robots.txt along with the time it was last checked, and URL Inspection will tell you whether Googlebot is allowed on a specific address. Expect recrawling and reindexing to take days to weeks; nothing will show on the day you deploy, and that is not a fault.

What the report says about it

Finding description

Your robots.txt denies [token] the whole site; the rule group for User-agent: [matchedAgent] contains Disallow: /. Googlebot is not an AI crawler and we do not track it as one; it matters here for a different reason. Google's own documentation states that a page can appear as a supporting link in AI Overviews or AI Mode only if it is indexed and eligible to be shown in Google Search with a snippet. A page disallowed in robots.txt reaches the index at most as a bare URL with no content, and Google will not show a snippet for it. So this is not a matter of worse odds: the condition Google itself sets is not met.

Recommendation

First check whether the block is intentional. On a staging or client preview site, on an intranet, and in any environment that is not meant to reach search at all, denying Googlebot everything is the correct setting and there is nothing to fix; just be aware that such a site will not appear in Google Search or in its AI answers, and that no other finding in this audit changes that. If the site is meant to be publicly discoverable, remove the block from robots.txt: either delete the Disallow: / line from the group that applies to Googlebot, or give Googlebot its own group with Allow: /. Watch out for one thing: a group naming Googlebot overrides the User-agent: * group as a whole, not just that one line. So if you create one, repeat inside it any partial rules (for example Disallow: /admin/) that should still apply. After deploying, verify the change in Google Search Console (Settings → robots.txt, or URL Inspection). Expect recrawling and reindexing to take days to weeks; nothing will change on the day you deploy.

Sources

Text verified 2026-09-14