JavaScript and AI crawlers: why content has to be in the HTML

A single-page app where content is filled in by JavaScript after load looks flawless in a browser. The question site owners rarely ask: does an AI crawler see the same thing? For most of them, the answer is no.

What the edge-log analysis showed

The best available source on this question is Vercel's analysis of its own edge logs, specifically nextjs.org plus two job boards on different frameworks, December 2024. Vercel has an advantage almost nobody else does: it can see directly what content a crawler downloaded and whether it executed JavaScript or not. It's the only independent source with a documented methodology we could find, and we couldn't locate a second one, which is worth saying plainly.

CrawlerOperatorRenders JavaScript?
GPTBotOpenAINo
ChatGPT-UserOpenAINo
OAI-SearchBotOpenAINo
ClaudeBotAnthropicNo
PerplexityBotPerplexityNo
BytespiderByteDanceNo
Meta-ExternalAgentMetaNo
Gemini (Google)GoogleYes
ApplebotAppleYes

Seven of the nine crawlers tracked download a page as raw HTML and never execute JavaScript at all, so they fetch the script files but don't run them. Content that only lands in the DOM after client-side hydration simply doesn't exist for them. Gemini and Applebot are the exceptions, and they render like a regular browser.

Why it matters

Indexability is, per Google's own documentation, a necessary condition for a page to appear as a supporting link in AI Overviews or AI Mode (see What documentably increases your chance of being cited by AI). Mind what does not follow from that, though: Googlebot does render JavaScript, so an SPA can be indexed by Google and can meet that condition. The problem lies elsewhere. A page whose main content is an empty <div id="root"></div> waiting for JavaScript is blank to the AI crawlers that do not render (GPTBot, ClaudeBot, PerplexityBot and the others in the table below), even though it looks complete in a browser.

It's not just text. Heading structure, tables and lists, that is everything a crawler uses to understand the hierarchy of a page's content needs to be in the HTML that arrives from the server, not added after the page has already loaded.

How to test exactly what a crawler sees

You don't have to guess or take a framework's word for it, because this is directly verifiable, and the process works on any site:

  • Fetch the page without running JavaScript with curl -A "GPTBot" https://your-site.com/page (or any tool that doesn't render), then check whether the page's main text is actually present in the returned HTML.
  • Open the same page in a browser and compare how much text you see there against what curl returned. Missing headings, paragraphs, or entire sections are exactly what render-less crawlers never see.
  • If there is a gap, check whether it covers content that matters (main text, headings, key facts) or just interactive elements (forms, filters, widgets), because a gap in the latter doesn't matter.
  • Test more than the homepage, because SPA frameworks often differ in how much content lands in the initial HTML on a landing page versus an inner page (a product detail, an article).

Which frameworks have a problem by default

The risk isn't the same across frameworks. Purely client-rendered apps (typically built as a single-page app with no additional server configuration) send the browser nearly empty HTML and fill in all the content with JavaScript afterward, and that is exactly the pattern render-less crawlers can't see. Frameworks that support server-side rendering or static generation by default, or as an easy option (Next.js, Nuxt, SvelteKit, Astro), send finished HTML in the very first response, regardless of whether the page is then "hydrated" with JavaScript on the client for interactivity.

Prerendering/dynamic-rendering services also exist, serving bots a prerendered version of a page while people get the regular SPA. It's a workable patch for a site where rewriting the whole architecture isn't realistic, but it adds another system to maintain and test, so for a new project, choosing SSR or static generation from the start is simpler and more reliable.

What to do about it

Server-side rendering (SSR) or static generation is the most reliable path, because content is present in the initial HTML response regardless of whether a crawler runs JavaScript. Next.js, Nuxt and similar frameworks support this natively, with no special extra configuration required.

A hybrid approach is a reasonable compromise where rewriting the whole app isn't realistic: statically or server-rendered core content (text, headings, key data) with JavaScript layered on top only for interactivity, that is forms, filters, animation. This works for both types of crawler at once, without compromising the user experience.

Verifying what a crawler actually receives belongs in ongoing site maintenance, not just a one-time launch audit, and the test above is worth repeating after any major frontend change, since a framework upgrade or a new library can quietly shift that ratio for the worse.

What we don't know

Vercel's methodology is limited to three specific sites over one month, so it's the best available source, not a definitive study of the whole web. Crawler behavior can also change faster than documentation keeps up; the last-checked date for each source is listed under Sources below.

There's also an open question about how quickly this could change. AI crawler behavior today tracks basic economics: rendering JavaScript is computationally more expensive than fetching raw HTML, and that cost difference multiplies across billions of crawled pages. There's no particular reason to expect that to change just because it would be more convenient for site owners, and it's more reasonable to expect the gap between rendering and non-rendering crawlers to persist.

If you manage a site with a frontend team, this question is worth putting directly into the definition of done for every new page, because it's cheaper to check during development than to track down two months later why a page never shows up anywhere.

Torumata's audit measures exactly this (check ACC-03): it compares the text a plain HTTP request receives against the text after a full browser render, and flags a difference greater than 40% of words as a finding. Run a free audit to see how much of your content AI crawlers actually see.

Sources