ACC-03

Main content requires JavaScript

CriticalCrawler access

What the check measures

Every sampled page is fetched twice. First as a plain HTTP request, with the text taken straight out of the delivered HTML, exactly what a crawler that runs no JavaScript sees. Then in a real browser (Playwright), which executes the scripts and lets the page settle. The word counts of the two versions are then compared.

The finding is raised when the difference exceeds 40 % of the rendered version's words. That threshold is not a guess buried in code: it lives in config/thresholds.json under rule_engine.ACC-03 as cheerio_vs_playwright_word_diff_ratio: 0.4, and it comes from the product specification. The formula is |words with JS − words without JS| / words with JS.

What the check does not do:

  • It does not run on every page. Only part of the site is sampled through a browser; it is the most expensive step of the whole audit. Pages outside the sample are not judged at all, so the absence of a finding does not mean the whole site is fine.
  • It does not say what is missing. It measures volume of text, not its content. A page missing 39 % of its text gets no finding, even if those 39 % held the entire price list.
  • It cannot tell the direction. When JavaScript removes text instead (a cookie banner in the static HTML that a script replaces with content), the difference comes out just as large; the check uses an absolute value.
  • It measures neither speed nor script errors. A page that renders correctly in a browser but takes eight seconds is fine as far as ACC-03 is concerned.
  • It does not judge whether a page is “important”. A finding on your privacy policy weighs the same as one on a product page.

The finding is page-level, so it appears once per affected page, grouped under a single code. Severity is critical.

How strong the evidence is

Documented effect

There is a measurement with a described methodology that actually found the effect. Not a guarantee of results, but more than an expert guess.

There is a measurement behind this check, not a hunch. Vercel analysed edge logs for three sites built on different frameworks (nextjs.org among them) and watched what AI crawlers actually do once they have fetched a page.

The result: GPTBot, ChatGPT-User, OAI-SearchBot, ClaudeBot, PerplexityBot, Bytespider and Meta-ExternalAgent do not execute JavaScript. They download .js files but never evaluate them. Whatever is not in the delivered HTML does not exist for them.

Two exceptions it is only fair to name: Gemini and Applebot do render. If those are your sole target, your situation differs. For most others it does not.

And one admission about the strength of the evidence: this is the only independent source with a described methodology that we were able to find. We looked for a second and did not find one. The data is from December 2024, so behaviour may have changed since, though none of the operators declares it either way. Hence the class “documented effect” rather than “necessary condition”: the measurement exists and is credible, but it is not backed by an operator's own documentation the way ACC-01 or ACC-07 is.

The practical consequence is nonetheless hard: if your main content is generated in the browser, this is not a matter of ranking lower. That content is simply not available to a model that came to your site.

How to fix it

Start by checking what a crawler actually receives. You need no tooling; a command line and a comparison with what you see in the browser will do:

curl -s https://your-domain/page | wc -w

If that returns a few dozen words while the page shows thousands in a browser, you have your answer. View source (Ctrl+U, not the inspector, which shows the state after scripts have run) and you will see whether the text is there at all or just an empty <div id="root">.

The fix comes in three levels, depending on how much work you can afford:

  1. Static generation (SSG): content that does not change per request (articles, documentation, product copy) should be rendered into HTML at build time. Cheapest to run, and the crawler gets everything.
  2. Server-side rendering (SSR): pages that depend on the request should be rendered on the server and sent as finished HTML. In Next.js, Nuxt and SvelteKit this is the default mode; the problem usually starts when a component is marked as client-side.
  3. A partial fix: when a site-wide change is out of the question, apply SSR at least to pages that carry content. A landing page without SSR costs you nothing; an empty product page costs you plenty.

Three mistakes people hit:

  • “We have SSR” does not mean this page has it. Modern frameworks pick the mode per component, and one client component wrapping the content is enough to strip it from the HTML.
  • Content loaded only after an interaction (accordions, “load more”, tab panels) must still be in the HTML even when visually hidden. Hiding it with CSS is fine; generating it with a script is not.
  • Prerendering for crawlers based on User-Agent is a dead end. Serving crawlers something different from people is risky and only covers the bots you name. Fix rendering for everyone, not for a list.

What the report says about it

Finding description

The text obtained statically (Cheerio render) differs from the text after JavaScript execution (Playwright render) by more than 40% of words. This means the main content is generated in the browser, and AI crawlers that don't execute JavaScript (most don't) never see it.

Recommendation

Implement server-side rendering (SSR) or static generation for the main content (e.g. Next.js SSR/SSG instead of a purely client-rendered page) so the text is present directly in the HTML response. If SSR isn't feasible site-wide, apply it at least to content-value pages (articles, products), not just the landing page.

Sources

Text verified 2026-09-12