Cloudflare changes its AI crawler defaults on September 15

Cloudflare sits in front of a large share of the web, and on 15 September 2026 it changes the defaults that decide whether AI bots get through. For two months the headlines have said “Cloudflare is going to block AI crawlers”. That shorthand overstates the change in one direction and understates it in another. What follows is based on Cloudflare's and Google's own texts (see Sources, verified 9 September 2026), and ends with a check you can run in two minutes.

Three categories instead of one switch

Since 1 July 2026 Cloudflare no longer sorts bots into “AI” and “not AI”. It sorts them by what they do with your content. There are three categories, each configurable on its own, including on the free plan:

  • Search: a bot that indexes your content so it can answer questions about it later. For this category Cloudflare says the site owner should expect referral traffic or other equitable compensation in return.
  • Agent: automation acting in real time on behalf of a specific person. This covers chat fetch bots (ChatGPT-User) and browser-driving agents.
  • Training: a crawler taking your content to train or fine-tune a model.

Each category offers three options: do not block, block only on pages that display ads, or block everywhere. What changes on 15 September is how those options are preset, and two separate things change, which secondhand coverage tends to merge into one.

Change one: new defaults for newly onboarded domains

The first change concerns default values. Cloudflare puts it like this:

For all new domains onboarding to Cloudflare, the categories of Training and Agent will be blocked by default on the pages that display ads, while Search will remain allowed by default.Cloudflare, Your site, your rules (1 July 2026)

Two qualifiers do the heavy lifting here, and both got lost on the way into the headline: new domains, and only pages that display ads. Search stays allowed. A site without ads is not changed by this default, and neither, on Cloudflare's own wording, is an existing zone.

An honest note about sources: some coverage states that the new defaults will also hit existing sites on the free plan. That is not what Cloudflare's own texts say: neither the blog post nor the documentation; both talk about “new domains onboarding to Cloudflare”. So we don't repeat a claim we could not confirm in the primary source. The check at the end of this article answers the question for your specific site either way.

Change two: this one reaches sites that have been on Cloudflare for years

The second change is quieter, and it applies to settings you already have. From 15 September, a bot that does several jobs at once is judged by all of its behaviours, and the most restrictive matching rule wins. The documentation for the original switch says so directly:

Mixed-purpose crawlers that combine Search and Training will also be blocked by all configurations to block AI training, including the legacy “Block AI bots” option.Cloudflare Docs, Block AI bots

In practice: if you ever flipped the one-click Block AI bots switch, mixed-purpose bots were exempt from it until now; the documentation described that as “This option excludes mixed-purpose bots that are used both for Training and for Search.” After 15 September they will not be exempt. Cloudflare names Googlebot, Applebot and BingBot specifically.

One distinction worth keeping straight: putting Googlebot in the Training bucket is Cloudflare's classification, not a statement from Google. Google has no separate training crawler. Google-Extended is a robots.txt token with no bot of its own, and it does not govern whether or how Googlebot crawls your site. We cover the difference between an opt-out token and a real crawler in AI crawlers and robots.txt.

Why a blocked Googlebot costs more than it looks

If you do have training blocked, work out the other side of the equation. Google's own documentation states:

To be eligible to be shown as a supporting link in AI Overviews or AI Mode, a page must be indexed and eligible to be shown in Google Search with a snippet. There are no additional requirements.Google Search Central, AI features and your website

Being indexed is therefore a necessary condition, not one optimisation among many. Being unreachable for Googlebot doesn't only cost you classic search rankings; it also removes your eligibility to appear as a cited source in AI Overviews and AI Mode. This is one of the few claims about AI visibility that rests on the operator's own documentation rather than on an estimate or a correlation.

And some proportion, because there is no point in overstating this: most sites will not be affected on 15 September at all. The ones that will are those with training blocking switched on, often still from the wave around 2025, when the switch was one click away and nobody has looked at it since.

The two-minute check on your own site

Whether this affects you is answered inside your own account. There is no point guessing when the answer is a few clicks away:

  1. Sign in to the Cloudflare dashboard and pick the domain (zone) in question.
  2. Open Security → Settings. Cloudflare's changelog links straight to dash.cloudflare.com/?to=/:account/:zone/security/settings.
  3. Find the card with the AI traffic settings. You will see either the older Block AI bots switch or the new trio of Search / Agent / Training.
  4. For each category, read which of the three options is selected: do not block, block only on pages with ads, or block everywhere.
  5. If training blocking is on, whether through the old switch or the Training category, decide deliberately whether that is what you want. From 15 September it will also apply to Googlebot, Applebot and BingBot.
  6. If you want to decline the new defaults, Cloudflare says you can mark that in the same Security settings any time before 15 September.

If you don't use Cloudflare, this change isn't yours to handle, but confirm nobody else runs a CDN in front of your site. On sites inherited from an agency or host, “we're not on Cloudflare” is wrong surprisingly often; a server: cloudflare header or the DNS records will tell you.

The second check belongs in Google Search Console: watch whether your pages are still being crawled and indexed. A CDN-level block doesn't show up as an error; the site keeps working for people, because the block targets bots. It shows up as a drop in crawling, a far slower and less obvious signal than a broken page.

What cannot be verified from the outside

Here is a limit we would rather state plainly. Whether a given site lets a given AI crawler through cannot be reliably determined from the outside. The server decides based on the visitor's user agent and IP address, so the only way to force an answer is to send a request pretending to be GPTBot or Googlebot. We don't do that. Our crawler identifies itself with its own user agent and never impersonates somebody else's bot.

What is visible from outside is circumstantial: a 403 from the edge, an interstitial, CDN headers. Those are indications, not proof, and a tool that turns them into certainty is overstating what it knows. The definitive answer is the setting in your account, per the procedure above.

What Torumata checks

From an audit's point of view, a CDN-level block is one more layer of the same problem we have worked on from the start: everything standing between your content and the bot meant to read it. The audit already compares robots.txt against a current list of AI crawlers, checks whether pages are indexable, and verifies the content is readable without running JavaScript.

The CDN setting is a fourth place where the same thing gets decided, and the only one the site owner does not hold in their own files. Nobody runs into it during an ordinary content review; it usually surfaces weeks later as a dip in Search Console. That is why we added two checks to the bot accessibility group, each resting on a different strength of evidence.

ACC-08 (warning) reports a Content-Signal line in robots.txt where ai-train or ai-input set to no, which is hard evidence from the file itself, flagged as a preference, not an enforced block. ACC-09 (informational) only detects from the response headers that the site sits behind a CDN whose defaults are changing, and sends you to verify it, because those settings live at account level and leave no trace, so the check says outright that it cannot confirm the real state from outside. Less than a customer would like, but true.

The change itself is not a catastrophe. It is a shift in default values that you are free to disagree with and reconfigure. What makes it awkward is only that it touches a switch nobody revisits for years, and that it takes effect quietly. Two minutes in your dashboard settle it.

Want to know what else stands between your content and the bots reading it today? Run a free audit.

Sources