SEO-08

Noindex on a valuable page

WarningCrawler access

What the check measures

The finding goes to a page that carries noindex (or none) and at the same time looks like it was meant to be indexed. “Looks like” is measurable here: at least 150 words of text and at least one internal link pointing at it.

Why it exists alongside ACC-07, which also reports noindex: ACC-07 reports every page with noindex, because for a language model it is invisible regardless of intent. SEO-08 is narrower and targets only where it is probably a mistake: the page has content and something links to it, so it likely was not meant to vanish from search. A login page or a basket will not meet the condition and gets no finding.

The 150-word threshold is repeated from the core set (where SIT-02 holds it in config/thresholds.json), not imported: the seo module is deliberately independent of the rule engine's configuration. So it is the same value in two places.

What the check does not do: it does not read X-Robots-Tag in HTTP headers (ACC-07 does), it does not count inbound links from other sites, and it cannot see noindex set by JavaScript. It also does not find out whether the page really is out of the index of Google or Seznam; that is shown by URL Inspection in Google Search Console and, for Czech sites, by Seznam Webmaster. The finding is page-level; severity is warning.

How strong the evidence is

Necessary condition

Without this, AI search has no way to show you at all. Google states it in its own documentation: a page must be indexed and eligible to be shown with a snippet. The strongest class we have.

This is one of two paid-module codes that concern a necessary condition, and it is the same one as ACC-07. A page with noindex is not in Google's index, and without the index it will not appear in AI features:

To be eligible to be shown as a supporting link in AI Overviews or AI Mode, a page must be indexed and eligible to be shown in Google Search with a snippet. There are no additional requirements.Google Search Central, AI features

Where this finding adds value over ACC-07: it does not tell you “this page is out of the index”, it tells you “this page is out of the index and probably nobody meant it to be”. That is the difference between a list you must read in full and a list worth your attention.

The commonest cause is still the same and bears repeating: a blanket noindex carried from staging into production. The site then disappears from search for months and nobody knows why, because nothing broke; nothing is simply shown.

How to fix it

For each reported page answer one question: should this page be in search? Because the check has already filtered out short and unlinked pages, the answer here is yes more often than with ACC-07.

If yes, find the source of the ban. It is in one of three places:

  1. In the HTML: <meta name="robots" content="noindex">. Look at view-source, not the inspector.
  2. In a system setting: WordPress has a “Discourage search engines” switch, other platforms an equivalent. It is usually left on from before launch.
  3. In one section's template: when every article of one category is in the list, it is one template, not fifty pages.

Remove the tag; do not replace it with content="index, follow": indexing is the default and an explicit permission adds nothing.

When the answer is no (an internal overview, a long thank-you page), ignore the finding, and consider whether such a page should be linked from the public part of the site at all. That link is what put it on this list.

What the report says about it

Finding description

The page has `noindex`, yet has substantial content and is linked from elsewhere on the site; it looks like a page you would want indexed, not one deliberately hidden. This is a standard recommendation for classic search engines (Google/Bing); we have no evidence it affects visibility in AI assistant answers.

Recommendation

Check whether `noindex` was intentional. If not, remove the directive from `<meta name="robots">`/the `X-Robots-Tag` header. We did not verify that the page really is out of the index: we only read the directive in the HTML. The real state is shown by a `site:` query in the search engine, by URL Inspection in Google Search Console and, for Czech sites, by Seznam Webmaster.

Sources

Text verified 2026-09-12