ACC-08

robots.txt sets a Content-Signal against AI training/input

WarningCrawler access

What the check measures

In robots.txt (the same file the audit fetches for ACC-01) we look for a Content-Signal: line and, inside it, the tokens ai-train and ai-input set to a blocking value. It typically looks like this:

Content-Signal: search=yes, ai-train=no, ai-input=no

The finding is raised when at least one of those two tokens is set to no. The report prints the actual lines and the actual tokens found, not a generic statement. The field names, values and tokens are not in code; they live in config/cdn-ai-defaults.json under content_signal, so they can be changed without touching the program.

This is purely a textual analysis of your own file. Nothing is inferred, no extra request is sent, and under no circumstances do we impersonate somebody else's crawler to check whether blocking works.

What the check does not do:

  • It ignores the search token. That one governs classic search; ACC-08 is only about AI training and AI input.
  • It cannot tell who wrote the line. It may be your decision, or your CDN may have added it for you through a managed robots.txt, which is precisely the case this finding exists for.
  • It does not verify that anyone honours it. That cannot be verified at all; see below.
  • It has nothing to do with whether your site runs behind a CDN. Anyone can type that line into the file by hand. Risk arising from a CDN's default settings is ACC-09, a different matter with a different kind of evidence.

The finding is site-level; severity is warning, not critical. The reason is in the next section and is worth reading.

How strong the evidence is

Effect not demonstrated

We recommend it because it does no harm or has some other benefit, but we promise nothing about whether it makes language models cite you. Nobody has demonstrated that yet.

Content Signals Policy is a convention introduced by Cloudflare in 2025: alongside rules about where a crawler may go, it adds to robots.txt a declaration of what the content may be used for, separately for search, for model training, and for insertion into a real-time answer.

How well it works is not documented, and Cloudflare says so itself. Its own documentation states that this is a signal, not a scraper blocker. No operator of a language model declares that it reads or honours Content Signals. That does not make the line useless: it is a clearly expressed wish from the owner of the content, which can carry legal and reputational weight. It only means that it enforces nothing technically.

That is exactly why the severity is only warning, while ACC-01 is critical. Disallow: / genuinely prevents a standards-abiding crawler from fetching the page. Content-Signal: ai-train=no is a request.

So what is the finding for? For letting you know the declaration is in your file. There are two common cases:

  • You wrote it and you know why. Then there is nothing to do; the finding is a statement, not a reproach. Take it as a reminder that you made this choice, and that it also covers tools that might otherwise have cited you.
  • Your CDN added it for you. A managed robots.txt can insert Content Signals on your behalf, and it takes one switch in a dashboard you may never have opened. You then expressed a wish you never uttered.

One more distinction that is easily blurred: ai-train and ai-input are not the same thing. The first says “do not train on my content”; the second says “do not put my content into the answer you are composing right now”. The second is precisely the case where an assistant would have cited you. Anyone who wants to stay out of training datasets while still appearing in answers should set ai-train=no and ai-input=yes, and that difference matters more than it looks.

How to fix it

First find out whether the line belongs there at all. Open https://your-domain/robots.txt in a browser and look at what is actually served, not at what is in your repository.

If you did not write the declaration, look for it in your CDN dashboard under managed robots.txt or in the section for managing AI bots. It is usually a single switch.

Then decide based on what you want from your content, and note that there are three options, not two:

What you wantHow to write it
To be in search and in AI answers, just not in trainingContent-Signal: search=yes, ai-train=no, ai-input=yes
To be in search, but otherwise keep AI at arm's lengthContent-Signal: search=yes, ai-train=no, ai-input=no
No restrictionsremove the Content-Signal: line entirely

What to watch out for:

  • ai-input=no is the consequential decision, not ai-train=no. The first costs you mentions in answers; the second only a place in training data. If you only meant the second, fix it.
  • The line does not belong inside a User-agent group. Write it on its own line, like Sitemap:.
  • If you did not want the declaration and you delete it, check the CDN switch too. Otherwise a managed robots.txt will put it back on the next rebuild.
  • Do not expect it to stop anything by itself. Anyone who wants training genuinely prevented needs blocking at the server or CDN level; the declaration is a statement of intent, not a lock.

What the report says about it

Finding description

robots.txt contains a Content-Signal line that explicitly sets: [the list of specific tokens] = no. Exact line found: “[the actual Content-Signal lines from your robots.txt]”. Content-Signal is a new, machine-readable preference right in robots.txt (not an instruction like Disallow). Per Cloudflare's own documentation it's a signal, not a technically enforced block: a crawler may or may not honor it. Still, it's a clearly stated intent by the site owner (or their CDN) about an ENTIRE CATEGORY of AI use of the content, not about one named bot the way ACC-01 is.

Recommendation

If you want AI systems to be able to use the content for training or for real-time answers, change the value of the tracked tokens (ai-train, ai-input) to “yes”, or remove the Content-Signal line from robots.txt entirely; its absence doesn't mean a ban, just no stated preference. If you didn't add this line yourself, it may come from your CDN or hosting (for example Cloudflare's managed robots.txt), so check the setting there.

Sources

Text verified 2026-09-12