ACC-05

Content likely available only after login or payment

WarningCrawler access

What the check measures

The check looks for two things on the same page: a login form (the HTML contains an input of type password) and very little publicly visible text. When a page has both, the finding is raised. The finding text tells you how many words we counted on that page, next to the threshold we compared it against.

“Very little text” means fewer than 100 words. It lives in config/thresholds.json under rule_engine.ACC-05 as max_words_for_short_text: 100, and it is explicitly marked there as "_status": "PROPOSAL", our own proposal pending revision against real data, not a number from the specification. The product spec gives only the heuristic “short text + login form”; the actual threshold is ours.

What the check does not do:

  • It does not flag the login page itself. Addresses such as /login/, /register/, /password-reset/ or /logout/ get no finding even when they meet the heuristic. A login page hides no content; it is the login. The list of patterns lives in config/thresholds.json under rule_engine.ACC-05. Until 19 September 2026 it did flag them, and that was this check’s most common false alarm.
  • It does not flag a variant address whose canonical points elsewhere. /article/?utm_source=mail gets no finding; /article/ does.
  • It cannot see a paywall without a password field. A wall that offers a “Subscribe” button and sends you to a payment page passes with no finding, even though that is exactly the case the finding is about.
  • It does not log in. We neither know nor can know what is behind the wall; the finding talks about what is in front of it.
  • It ignores a login box that comes from the template. A form with a password field only counts when it sits in the page's main content. A box in the header, footer or sidebar that the template puts on every page of the site does not produce a finding - it says nothing about content being behind a login. That was not the case until 20 September 2026, and on an e-shop it produced nine findings on public product listings, critical at the time.
  • It does not judge the value of the content. A short page with a login form in its content gets the finding just like a locked article.
  • It does not check for paywalledContent structured data, which is the proper way to tell Google that content is paid.
  • It never runs on pages the crawler did not visit.

The finding is page-level; severity is warning. It was critical until 22 September 2026; the owner lowered it because the check rests on a heuristic (little text plus a password form), and critical is reserved in this product for what we detect with certainty.

How strong the evidence is

Necessary condition

Without this, AI search has no way to show you at all. Google states it in its own documentation: a page must be indexed and eligible to be shown with a snippet. The strongest class we have.

A crawler will not log in and will not pay. What is behind the wall is not in the index, and without the index the page does not appear in AI answers. Google states it in its own documentation, and Microsoft Copilot builds on Bing's index, so the same holds:

To be eligible to be shown as a supporting link in AI Overviews or AI Mode, a page must be indexed and eligible to be shown in Google Search with a snippet. There are no additional requirements.Google Search Central, AI features

What does not follow from that: that you ought to unlock the content. Paid content is a legitimate business model, and this finding is not a call to abandon it. It is a statement of price: pages behind a wall do not exist for language models, and therefore neither do they exist for the answers that might have cited you.

The firm part is the first half (what a crawler cannot see is not indexed). The second half (that pushing an excerpt in front of the wall leads to citations) is not documented and we do not promise it. All we know is that without it the chance is zero.

One more piece of honesty about the detection itself: it is a heuristic, not a measurement. The 100-word threshold is our estimate, marked in the configuration as a proposal pending revision, and our crawler cannot by definition see what is behind a login - so we cannot tell a page with hidden content apart from a page that hides nothing and simply has little text. Login and registration pages are excluded from the finding (since 19 September 2026), but before you change anything, look at which page it actually is.

How to fix it

First decide which of three groups the page belongs to, because each is handled differently:

  1. A login to an admin area or a customer account. Nothing to fix. Quite the opposite: put such a page under Disallow in robots.txt and mark it <meta name="robots" content="noindex">. It has nothing to offer an answer.
  2. Paid content you intend to keep selling. Keep the wall and push in front of it as much as you can bear.
  3. Content restricted by accident: a section locked long ago for a reason that no longer exists. Unlock it.

For the second group, the question is what exactly to leave outside. A simple measure works: leave out enough for a reader (and a model) to grasp what is being claimed inside, but not enough to substitute for the subscription.

  • The headline and the standfirst: in full, not truncated with an ellipsis mid-sentence.
  • The first two or three paragraphs, enough to make the subject clear.
  • The key findings or an abstract: for specialist content this is often enough, and it is the most frequently cited part of all.
  • The date, the author and the sources: these have no business behind the wall; they are machine-readable context that AUTH-02 and AUTH-03 measure.

One technical piece belongs with it: paid content can be declared as paid, properly. In structured data the locked section is marked isAccessibleForFree: false with a cssSelector pointing at it. Google's documentation says this markup is what lets it tell paywalled content apart from cloaking, which violates its spam policies, and that is the difference between doing this correctly and serving crawlers something other than people.

{
  "@context": "https://schema.org",
  "@type": "NewsArticle",
  "isAccessibleForFree": false,
  "hasPart": {
    "@type": "WebPageElement",
    "isAccessibleForFree": false,
    "cssSelector": ".paid-content"
  }
}

A final warning, because it is a trap you can walk into with good intentions: do not serve crawlers the full text and people a wall. That breaks Google's policies, and the markup above is the sanctioned way around it, not User-Agent detection.

What the report says about it

Finding description

On this page we counted [wordCount] words of publicly visible text (our heuristic threshold is [threshold] words) together with a form containing a password field in the page's main content. That pair usually means the main content only becomes available after logging in or paying; AI crawlers don't log in or pay, so whatever stays behind the login does not exist for them. This is a heuristic, not a measurement: a crawler cannot by definition see what is behind a login, so we cannot tell a page with hidden content apart from a page that hides nothing and simply has little text. Login, registration and sign-out pages and password-reset forms are excluded from this finding - they hide no content, they are the login. A login box that the template places in the header, footer or sidebar of every page on the site does not produce a finding - it says nothing about content being behind a login. Variants of the same address that differ only by a query parameter produce one finding, not several.

Recommendation

Content that should be reachable without logging in must be published at least as a summary or an opening section in front of the login/paywall (e.g. the first few paragraphs, an abstract, key takeaways) - as an ordinary public page a crawler can fetch without signing in and without paying. This is a necessary condition, not an optional extra: whatever stays behind the login does not exist for AI systems. An `llms.txt` file is no substitute - it is only a list of links, and a crawler that hits a login behind them gets nothing (see ACC-02).

Sources

Text verified 2026-09-19