Structured data and AI: what's documented and what isn't

Structured data (JSON-LD, schema.org) is, along with llms.txt, one of the topics GEO advice, that is, advice on preparing a site for generative search, talks about most confidently. The primary-source reality is more careful: the benefit for classic search is well documented; for AI assistants, it isn't.

What Google says

Google's AI optimization guide explicitly warns against "overfocusing on structured data" and states that structured data is not required for generative AI search, because no special markup is needed for AI Overviews or AI Mode. Its structured-data intro page doesn't mention AI Overviews or AI Mode even once. In May 2025, John Mueller wrote in an article about AI search that structured data "makes pages eligible for certain search features and rich results", but without any claim about AI ranking.

OpenAI, Anthropic and Perplexity don't declare reading structured data at all. Perplexity even states directly that PerplexityBot "is not used to crawl content for AI foundation models", so even for the one bot that does visit sites, it's unclear what the structured data on those pages is actually used for.

The best available measurement

Ahrefs compared 1,885 pages that added JSON-LD themselves against 4,000 matched control pages (August 2025 – March 2026, difference-in-differences). Result: ChatGPT +2.2%, Google AI Mode +2.4%, and both are statistically insignificant. Google AI Overviews −4.6%, and that result is significant. The authors themselves wrote the difference "could easily be random noise."

The Otterly controlled experiment (March 2026)

The field's one controlled experiment found an increase for Google AI Overviews, but a decrease for ChatGPT, Gemini, Perplexity and Copilot. Information added only to the schema (not to the page's visible text) wasn't used by any of the tested platforms. That's the key takeaway: schema doesn't replace content, it only adds a machine-readable description on top of it.

The gap between Google AI Overviews (an increase) and every other platform (a decrease) has a plausible technical explanation, even though none of the sources states it outright: AI Overviews runs on the same index and the same page evaluation as classic Google Search, so structured data may play a role there through rich-result qualification rather than through the language model itself. ChatGPT, Perplexity, Gemini and Copilot assemble answers differently, namely from whatever their own crawler fetched and the language model judged relevant, where schema isn't a priority input.

Claims you shouldn't trust

  • "Microsoft confirmed schema helps Copilot" (attributed to Fabrice Canel, March 2025), yet no such statement could be found in Bing's blog on structured data.
  • "Generic schema is 18 percentage points worse than no schema": from a self-published, non-peer-reviewed preprint by the owner of a commercial GEO agency. The individual numbers check out, but the overall "has schema vs. no schema" effect was statistically insignificant (p = 0.296), and the cited comparison is missing both a p-value and a subgroup size.
  • "Ahrefs measured a 2.2% improvement": the number is real, but the authors themselves warn against reading it as evidence of an effect, and citing it alone omits the significant decrease for AI Overviews in the same study.
  • "An experiment proved models don't parse schema" (Mark Williams-Cook, February 2026): the experiment is real and verifiable, but it covers one page and one prompt; the author himself writes it doesn't prove models ignore schema entirely.

The common thread in all four claims: each takes one imperfect measurement and adds a level of certainty the measurement itself doesn't support. That's understandable, because a clean answer like "drives traffic" sells better than "the effect is within the range of noise", but it's exactly the kind of oversimplification this article is trying to avoid.

FAQPage: what Google dropped, and what it didn't

On 2026-05-07, Google discontinued FAQPage rich results, that is the expandable question-and-answer display directly under a search result. The Search Console report and Rich Results Test for FAQPage disappeared in June 2026, and API support followed in August 2026. The FAQPage type itself remains valid and Google still parses it for page understanding. Whether AI platforms read it too is something we do not know: Perplexity does not mention structured data in its own crawler documentation, OpenAI and Anthropic do not declare reading it, and for Bing the claim could not be sourced. What no longer holds is that FAQPage produces the expandable Q&A display in Google results, because that specific feature is gone.

Where to add structured data first

If you're deciding what to prioritize, follow what Google actually qualifies for rich results, not assumptions about AI: Organization/WebSite on the homepage (operator identity), Article/BlogPosting on articles (author, publication date, exactly the signals that establish freshness and trustworthiness), Product on product pages, and FAQPage/HowTo wherever the page already genuinely has a question-and-answer or step-by-step pattern. Don't build a FAQ section artificially just to justify the markup, because schema describes content that's already there, not content you'd like to have for markup's sake.

What this means for your site

You can honestly describe structured data as a machine-readable, unambiguous description of your content that helps Google qualify pages for rich results in classic search. What you can't honestly claim is that AI systems read it or that it increases citation rates in their answers. No major provider declares that, and the best available measurements show an effect within the range of noise, with a measured decline on some platforms.

Unlike llms.txt (see the dedicated article), structured data still counts as a genuine recommendation, not just a neutral option, because Google recommends it itself, for classic search, just not as a lever for AI visibility.

One last note on the data itself: schema should be an accurate, verifiable description of what's genuinely on the page, that is a publication date that's correct, an author who exists, a price that's current. Inaccurate or stale schema is worse than none at all, because search providers (Google explicitly) penalize a mismatch between structured data and a page's visible content.

Torumata generates JSON-LD snippets and flags missing ones in an audit (SCH-01–SCH-03), as an investment in classic search and future machine readability, not a promise of AI citation. Run a free audit to see where structured data is missing on your site.

Sources