How the score is calculated
The score is a weighted average of five categories and a quarter of it is scored by a language model. Here is the whole calculation: where its practical ceiling lies and why the number shifts a little between two audits.
The AI readability score is a single number from 0 to 100 summarising how ready your site is for AI search. It is a useful number and worth tracking. It is not a race to 100, though: 97 to 98 is the practical ceiling, and the difference between 94 and 96 tells you less than one unfixed critical finding does.
This page explains how the number is calculated, where its upper limit comes from, and why two audits of the same site can differ even when you changed nothing.
What the score is made of
It is a weighted average of five categories. Each is calculated separately on a 0 to 100 scale and enters the total with its own weight (the weights live in the audit configuration, not in code):
| Category | Weight | How it is produced |
|---|---|---|
| AI access | 30 % | Deterministic checks ACC-…: robots.txt, noindex, content behind a login or behind JavaScript. |
| Content for AI | 25 % | Scored by a language model. The only category with no checks of its own, see below. |
| Structure | 20 % | Deterministic checks STR-… and SIT-…: headings, descriptions, sitemap. |
| Schema | 15 % | Deterministic checks SCH-…: JSON-LD structured data. |
| Authority | 10 % | Three signals: links to sources, a named author, a last updated date (AUTH-01 to AUTH-03). |
Three categories (access, structure, schema) are calculated by subtraction: the category starts at 100 and every finding takes points off according to its severity and to how many pages it affects. An informational finding takes nothing off. A site where we found nothing in those categories scores 100 in them, and that is a common result.
Authority is not calculated by subtraction but from three weighted signals: links to sources (0.4), a visible author (0.3) and a last updated date (0.3). It is the one category whose effect on visibility in AI answers is backed by measurement (Princeton, KDD 2024: links to sources raised visibility by 28 %).
The other scores in the report (SEO, accessibility, colours, AI citations, design) are calculated separately and do not feed into the AI readability score. That is why the report shows them as separate numbers rather than one grade for the site.
Content for AI: the category a language model scores
A quarter of the total score comes from reading, not from measuring. The “content for AI” category has no checks of its own: the model reads every page in the sample and gives it two marks on a scale of 1 to 5.
- Does the page answer directly? Does it give concrete facts to the reader's main questions: figures, prices, deadlines, conditions or a procedure?
- Can the content be extracted? Is it broken up by subheadings, with concrete facts in lists and tables rather than buried in prose?
Each mark is averaged across the sampled pages, the two averages are averaged again, and the result is mapped linearly onto 0 to 100: a 1 is 0 points, a 5 is 100 points. Pages where the analysis failed are left out of the average rather than scored zero.
What a 5 means is best shown by the instructions the model is given:
| Mark | What it means for “does the page answer directly?” |
|---|---|
| 1 | Slogans, headings or links elsewhere, nothing more. |
| 2 | The page names what it offers but says nothing further about it. |
| 3 | It describes the offer or topic but gives no concrete facts. |
| 4 | It gives concrete facts for some of the main questions and leaves others unanswered. |
| 5 | It answers the main questions with concrete facts: figures, prices, deadlines, conditions or a procedure. |
This is the crux of it: only a page with concrete figures, prices or procedures gets a 5. A contact page, an “about us” page, a category listing or a marketing page will not get one by their very nature, even when they are well written and do exactly the job they are on the site for. And because the average covers the whole sample, pages for which a 4 is the right mark pull it down too.
Why 97 to 98 is the realistic ceiling
Just put the numbers in. A site where we found not a single finding and which has all three authority signals has four categories at 100, and that is 75 points out of a hundred (30 + 20 + 15 + 10). The remaining 25 points are handed out by content for AI:
| Average of the model's marks | Content for AI | Overall score (other categories at 100) |
|---|---|---|
| 3.5 | 62.5 | 91 |
| 4.0 | 75 | 94 |
| 4.3 | 82.5 | 96 |
| 4.5 | 87.5 | 97 |
| 4.75 | 93.75 | 98 |
| 5.0 | 100 | 100 |
So a 100 means every page in the sample got a 5 on both marks: that even the contact page answers with concrete figures and keeps them in a table. That does not happen on a real site and chasing it is not worth your time. We treat 90 and above as a very good result, which is grade A in the report.
An example from our own kitchen: our site torumata.com averaged 4.33 in its production audit, which is 83 for content for AI and 96 overall, even though the other four categories were at 100. Those four missing points are not a fault. They are there because a contact page will never answer with a price list.
Why the score differs between two audits
The score is a snapshot in time, not a constant. Two audits of the same site can differ by a few points without you changing anything. There are three reasons and all of them are a property, not a bug:
- The model is not deterministic. The same page gets a 4 one time and a 5 the next. One such shift on one page out of ten moves content for AI by about 1.25 points and the overall score by about 0.3 points; several of them add up to a few points.
- The sample of pages changes. The content analysis does not cover the whole site, only a sample. When pages appear or disappear in the meantime, the average is taken over a different set.
- The site itself changes. The crawler saw what was there at that moment: an outage, a temporary redirect or a migration to a new CMS all show up in the score.
The practical consequence: do not read a difference of up to two points as an improvement or a decline. Compare the trend across several audits, and above all the specific findings that appeared or went away.
The free audit's technical score is a different quantity
A free audit does not run the language model content analysis at all. Its score is therefore calculated from four categories only, with their weights rescaled to add up to a hundred percent again. It is called the technical score, not the overall score, and the report says so under the number.
That has one surprising consequence which we would rather state up front: the same site gets a higher number in a free audit than in a paid one. The category that takes points off most often is simply missing from it. It is not a trick and not a penalty for paying; they are two different quantities with two different names.
What to look at instead of chasing 100
In this order:
- High severity findings. A critical finding means AI either cannot reach your content or is not allowed to use it. One of those costs you more than ten points of score.
- The trend between audits. Is the number of findings falling? Have the critical ones gone? That answers the question “is what we are doing working”, which a single number cannot.
- The category that lags behind the rest. A score of 96 made of five hundreds and one eighty is a different situation from 96 made of five ninety-fives.
- Pages the model marked 1 or 2. That is where rewriting content really pays off: they lack the concrete facts an AI could cite.
And one closing sentence, because this is the whole point: a site scoring 93 that answers customer questions with concrete figures is better prepared for AI search than a site scoring 97 with nothing but perfectly filled in metadata.
Sources
- Google Search Central: AI features and your website (accessed 2026-09-19)
- Princeton: GEO: Generative Engine Optimization (arXiv 2311.09735, KDD 2024) (accessed 2026-09-19)