ACC-09

Site runs behind a CDN: check the AI bot block setting (it can catch Googlebot too)

InfoCrawler access

What the check measures

From the response headers of your home page we recognise whether the site runs behind a CDN whose default rules for AI crawlers are changing. For Cloudflare we detect it from a server: cloudflare header or the presence of cf-ray, cf-cache-status or cf-mitigated. No extra request is sent for this; the headers come from the crawl that happens anyway.

The finding is raised only when robots.txt carries no Content-Signal directive at all, neither blocking nor permitting. That suppression matters: the presence of such a directive means somebody has been at the AI-crawler settings, and a vague “you might want to look” would add nothing. The blocking variant is handled separately and deterministically by ACC-08. Without this rule, practically every site on Cloudflare would get the finding, and that is a very large share of the internet.

What the check does not do, and for ACC-09 this is the important part:

  • It does not claim your site blocks AI crawlers. Nor does it claim it does not. We do not know, and it cannot be determined from outside.
  • It does not impersonate somebody else's crawler. The only way to verify the blocking would be to pose as Googlebot or GPTBot and see whether the CDN lets us in. We do not do that; it is precisely the behaviour we criticise in our own reports.
  • It does not read your CDN account settings. We have no access to them and want none.
  • It recognises no CDN beyond those in config/cdn-ai-defaults.json. Adding another is a pure configuration change; no provider name is hard-coded anywhere.

That is why this is an informational finding, not a warning: it is a pointer to something worth checking, not a statement about the state of your site. It costs you no score at all: informational findings carry a weight of zero in our model.

How strong the evidence is

Necessary condition

Without this, AI search has no way to show you at all. Google states it in its own documentation: a page must be indexed and eligible to be shown with a snippet. The strongest class we have.

This badge needs explaining, because it only half fits. The consequence the finding warns about is a necessary condition, the best-documented thing we have. Whether that consequence applies to you is something we cannot measure. Hence the informational severity, and hence the “go and look” advice.

On 15 September 2026, two different mechanisms change at Cloudflare, and they are easily confused. We separate them deliberately, because press coverage merged them into one and added a scope the primary source never claims:

MechanismWho it affectsWhat changes
New default for new domainsExclusively domains onboarding to Cloudflare, and only their pages that display ads. Established domains and existing accounts are untouched.The Training and Agent categories will be blocked by default. Search stays allowed.
End of the exemption in the “Block AI bots” switchAnyone who has that switch turned on, possibly for years.Until now it exempted mixed-purpose bots. From now on it hits them too, Googlebot among them.

The second row is the serious one, which is why the finding leads with it. Googlebot out of the index is not merely a loss of classic search:

To be eligible to be shown as a supporting link in AI Overviews or AI Mode, a page must be indexed and eligible to be shown in Google Search with a snippet. There are no additional requirements.Google Search Central, AI features

Put plainly: a switch turned on because of “AI training” may, after 15 September 2026, knock a site out of the necessary condition for appearing in Google's AI features at all. That confusion is the reason this finding exists.

One last clarification worth stating: classifying Googlebot, Applebot and BingBot as “mixed-purpose” bots is Cloudflare's decision, not a statement by Google, Apple or Microsoft. Google has no dedicated training crawler of its own; Google-Extended is merely an opt-out token with no robot behind it.

How to fix it

This is five minutes of work and you need nobody else for it. In the Cloudflare dashboard go to Security → Settings. Both relevant items sit on that one page, not in two different places:

  1. “Block AI bots”: the older one-click switch. If it is on, this is the one that from 15 September 2026 also hits Googlebot. Decide whether that is what you want.
  2. “Configure AI bot policies”: the finer-grained setting by category (Search / Training / Agent). Here you can leave Search allowed and restrict training only.

Cloudflare itself published a direct dashboard link to that page, so there is nothing to hunt for.

When deciding, it helps to separate three things that get muddled:

  • Search: this includes Googlebot. Leave it on unless you want to vanish from search and from AI overviews at the same time.
  • Training: using your content to train models. A legitimate thing to block if it matters to you.
  • Agent: bots that fetch a page in response to a user's question right now. This is the category that produces a citation in an assistant's answer. Blocking it means giving those up.

And two notes on top:

  • If you change nothing, at least write down what is on. It will save an hour when somebody asks a year from now.
  • Do not confuse this with “AI Crawl Control”. That is a different dashboard section managing textual directives in robots.txt, and that is ACC-08, not this finding.

What the report says about it

Finding description

Your domain's response headers show the site running behind [your CDN's name]. As of [the effective date], [your CDN's name] is changing two separate things at once, both through account-level/firewall settings with no observable trace in robots.txt or in HTTP headers, so Torumata CANNOT verify either from outside: 1) Domains newly onboarding to [your CDN's name] get the Training and Agent bot categories blocked by default on pages that show ads (Search stays allowed). This does not apply to domains already on [your CDN's name]. 2) Anyone with the older one-click "Block AI Bots" toggle turned on (whenever they turned it on) has so far had mixed-purpose crawlers excluded from that block: Googlebot, Applebot, BingBot (a classification [your CDN's name] makes itself, not a statement from Google, Apple, or Microsoft). That exclusion ends on [the effective date], so even a long-standing block can newly catch Googlebot. That's more serious than losing AI training access: being in Google's index is a necessary condition for showing up in Google AI Overviews and AI Mode too.

Recommendation

If the site runs on Cloudflare, open the dashboard → your domain → Security → Settings and check two entries there: "Block AI bots" (the older toggle; after September 15, 2026 it can newly catch Googlebot, Applebot, and BingBot too) and "Configure AI bot policies" (the new Search/Training/Agent category controls, which apply to newly connected domains with pages that show ads). Both live on the same dashboard page (Cloudflare's own opt-out link: dash.cloudflare.com/?to=/:account/:zone/security/settings). It takes a few minutes and is the only reliable way to know the real state; it can't be verified from outside.

Sources

Text verified 2026-09-12