ACC-07

The page blocks indexing (noindex)

CriticalCrawler access

What the check measures

On every live page (a 2xx response) we read whether the page forbids its own indexing. The finding is raised when the page carries noindex or nofollow, whether in <meta name="robots">, in <meta name="googlebot">, through the value none (shorthand for both at once), or in the X-Robots-Tag HTTP header.

nofollow is reported even without noindex, and it is worth knowing why. Such a page can still be indexed, but the links on it lead nowhere: content it points at, with no other route in, stays undiscovered. It is a quieter defect than noindex, but similar in consequence.

What the check does not do:

  • It does not judge whether the ban is appropriate. A login page, a basket, a post-submit thank-you page or a print view are supposed to carry noindex. A finding on those is correct and should not be “fixed”.
  • It never runs on pages with an error status: for 4xx and 5xx, indexability is moot; that is ACC-04's business.
  • It does not catch noindex set by JavaScript at runtime. A crawler that runs no scripts will not see it either, but neither will we find it in the HTML.
  • It does not look at robots.txt (that is ACC-01) or at CDN-level blocking (ACC-09). Those are three independent routes out of the index, and each has its own code.
  • It does not find out whether the page is in the index. We only read the directive the page carries; Torumata cannot see into the index of Google or Seznam. The real state is shown by a site: query in the search engine, by URL Inspection in Google Search Console and, for Czech sites, by Seznam Webmaster.

The finding is page-level; severity is critical.

How strong the evidence is

Necessary condition

Without this, AI search has no way to show you at all. Google states it in its own documentation: a page must be indexed and eligible to be shown with a snippet. The strongest class we have.

Together with ACC-01 this is the hardest claim in the whole audit. A page carrying noindex is not in Google's index, and without the index it will not appear in AI Overviews or AI Mode:

To be eligible to be shown as a supporting link in AI Overviews or AI Mode, a page must be indexed and eligible to be shown in Google Search with a snippet. There are no additional requirements.Google Search Central, AI features

Unlike robots.txt, which forbids a crawler from fetching a page, noindex lets it fetch the page and tells it not to file it away. The difference is practical: a page disallowed in robots.txt occasionally shows up in the index as a bare URL with no content (Google knows of it from links), whereas noindex removes it reliably. That is exactly why noindex is the more effective way to hide a page, and the more dangerous one when it gets there by accident.

And by accident it often does. The most common cause we see in practice is a staging setting carried into production: the development environment has a blanket noindex so it does not get indexed, and the value travels along with the deployment. The site then disappears from search for months and nobody knows why.

Where we are careful: nofollow has weaker backing than noindex. That a nofollow link passes no signals is documented by Google; that a page behind such a link will never be discovered is not absolute; it can be found another way (from the sitemap, from another page, from outside). So treat nofollow as a warning about a narrowed path, not a verdict.

How to fix it

For each reported page, answer one question first: should this page be in search at all? For a login page, a basket, a thank-you page or an internal filter the answer is no, and the finding is just a statement of fact, not a fault.

When it should be and is not, hunt for the source of the ban. It can sit in three places and look identical from the outside:

  1. In the page HTML: <meta name="robots" content="noindex"> or content="none". Check view-source (Ctrl+U), not the inspector.
  2. In an HTTP header: X-Robots-Tag: noindex. You will not see it in the page source at all; only a request from outside reveals it.
  3. In a system setting: WordPress has a “Discourage search engines from indexing this site” switch; Shopify, Webflow and most frameworks have an equivalent. It is usually left on from before launch.

One command reveals both the header and the tag:

curl -sI https://your-domain/page | grep -i x-robots-tag
curl -s  https://your-domain/page | grep -i 'name="robots"'

When the page should be indexed, remove the tag entirely. Do not replace it with content="index, follow". Indexing and following links are the default behaviour; Google's own specification says of the value all that it is “the default value and has no effect if explicitly listed”, and index, follow is the same thing in other words. An explicit permission improves nothing and only raises the odds that someone gets confused next time.

Two final things that trip people up:

  • Disallowing a page in robots.txt and giving it noindex does not work. Google documents this explicitly: if the page is blocked by robots.txt, the crawler will never see the noindex rule, and the page can still appear in search results, for instance if other pages link to it. To get a page out of the index, leave it reachable and rely on noindex.
  • Recovery takes time. Removing the tag is not the same as an instant return to the index; the page has to be crawled again. Request indexing in Search Console and do not expect a result by tomorrow.

What the report says about it

Finding description

The page carries a `<meta name="robots">` tag with `noindex` or `nofollow`. Search engines and language models therefore won't treat it as a source they can answer from; its content is invisible to them, even though a visitor sees it normally.

Recommendation

If the page is meant to be found publicly, remove `noindex` from `<meta name="robots">`. Keep the block only on pages that should stay out of search: thank-you pages after a form, internal listing filters, staging versions. We only read the directive on the page: Torumata does not check whether Google or Seznam actually has it in the index. The real state is shown by a `site:` query in the search engine, by URL Inspection in Google Search Console and, for Czech sites, by Seznam Webmaster.

Sources

Text verified 2026-09-12