What documentably increases your chance of being cited by AI

Once you rule out llms.txt (see the dedicated article) and structured data (see Structured data and AI) as documented levers for citations, one question is left: what actually works? Here's the answer, ranked by strength of evidence, starting from the hardest fact down to what has no support at all.

1. Indexability as a necessary condition

"To be eligible to be shown as a supporting link in AI Overviews or AI Mode, a page must be indexed and eligible to be shown in Google Search with a snippet. There are no additional requirements."Google Search Central, AI features in Search

This is the hardest fact in the whole survey, because it comes straight from Google's own documentation. For a page to even be eligible as a supporting link in AI Overviews, the generated summary above the classic results, or in AI Mode, Google's conversational search mode, it must first be indexed and eligible for regular Search with a snippet. Google states there are no additional requirements. Microsoft Copilot is built on Bing's index, so the same logic applies there too.

The consequence: indexability checks, that is whether robots.txt disallows you, whether you have a sitemap.xml, whether a page returns 200 instead of an error, aren't "also nice for AI". They're the ticket to entry. Without it, nothing else matters.

A basic indexability checklist you can run yourself: the page isn't disallowed in robots.txt for Googlebot, has no noindex meta tag (an instruction telling a search engine to keep the page out of its index), returns HTTP 200 (not a redirect or an error), is listed in sitemap.xml, and has at least one internal link pointing to it from elsewhere on the site. Missing any one of these five means the page has no chance of appearing in AI Overviews or AI Mode, and this holds regardless of how good its content is.

2. AI crawlers don't render JavaScript

An analysis of Vercel's edge logs (nextjs.org and two job boards on different frameworks, December 2024) found that GPTBot, ChatGPT-User, OAI-SearchBot, ClaudeBot, PerplexityBot, Bytespider and Meta-ExternalAgent don't execute JavaScript, so they download the files but never run them. Gemini and Applebot are the exceptions, and they render. This is the only independent source we could find with a documented methodology; we couldn't locate a second one.

In practice: content that only appears in the page after client-side JavaScript runs (typical of React or Vue apps without server-side rendering, that is, without the content already being rendered on the server) is invisible to most AI crawlers. Covered in detail in JavaScript and AI crawlers.

3. Content citability and the field's one peer-reviewed experiment

Princeton's GEO paper (arXiv 2311.09735, published at KDD 2024) is the only peer-reviewed experiment in the entire GEO and AEO field, that is, in preparing a site for generative search and answer assistants. The authors ran 10,000 queries and measured from the outside, with no access to the internals of the systems tested (a black-box test). What measurably increased a page's visibility in generated answers:

FactorMeasured visibility increase
In-text citations+41%
Concrete statistics+33%
Links to sources+28%

It's the best-documented lever the field has, and it's about content, not technical setup, that is writing, not configuration. Links to sources are exactly what Torumata's AUTH-01 check measures; citations and concrete statistics aren't scored in the audit today, but the same logic applies: text with verifiable numbers and links to primary sources has a measurably higher chance of being cited.

What this looks like in practice: the sentence "our tool saves time" carries none of what the experiment measured. The sentence "a 2026 internal study across 240 customers cut average processing time by 31% (source: [link to the study])" carries a concrete statistic and a source link at once, which is exactly the combination that increased citation rates the most in the Princeton experiment.

4. The link to classic search ranking is weak, and weakening

Ranking well in classic search helps, but it's far from a guarantee of an AI citation. Ahrefs found that the share of AI Overview citations coming from the classic top 10 fell from roughly 76% (July 2025) to 37.9% (March 2026, 863,000 queries). Semrush, on a sample of 200,000 keywords, measured an overlap of only around 20–26%, and just 46% of desktop AI Overviews link to the first organic result.

In short: a good ranking is necessary per point 1 (you need to be indexed and eligible), but far from sufficient for an AI citation. Practical consequence: a page sitting at position five or ten in classic search can end up cited in an AI answer more often than the page at position one, if it has a better combination of citations, statistics and source links from point 3. That's counterintuitive from a traditional SEO angle, but it matches how generative search actually assembles an answer, because it doesn't pick one "best" result, it blends passages from several sources.

5. What has no support

Domain authority, backlinks and domain age: these are classic SEO claims carried over into GEO without their own evidence. The major tools tracking AI citations don't analyze these factors at all. We don't build them into scoring, and we wouldn't recommend betting a strategy on them until that changes.

What to do about it

The order worth working through mirrors the strength of evidence above: fix indexability first (point 1), then confirm AI crawlers actually see the page's main content rather than an empty shell waiting on JavaScript (point 2), and only then get into writing, which means adding citations, concrete numbers and source links (point 3). Investing time in point 3 before points 1 or 2 are sorted doesn't pay off: even beautifully written text a crawler never receives is invisible as far as citations are concerned.

What we couldn't verify

To be honest: we couldn't find Microsoft/Bing's position on llms.txt, we couldn't locate a second independent source on JavaScript rendering, and we neither measured nor found elsewhere a correlation between domain age and citation rate. These aren't details we're hiding; they're the current limit of what's verifiable.

Curious how your own indexability and JavaScript rendering hold up? Run a free audit.

Sources