Recommendations from a language model

LLM- codes came from a reading, not a measurement: a model read the text of your pages. Here is how they arise, why they never get critical severity, and how to read them.

Codes starting with LLM- differ from the rest of the report in one fundamental way: they did not come from a measurement, they came from a reading. Every other finding was produced by a deterministic check: we fetched a file, counted characters, compared numbers. These arose because a language model read the text of your pages and wrote down what it thought could be improved.

That is a different kind of claim and it deserves different handling. This page explains exactly how they come about, why they carry different weight from the checks beside them, and what to do with them.

How a recommendation comes about

In two steps. First a model reads the individual pages of the sample and extracts a structured summary from each. The sample is at most 10 pages on a free audit and 30 on a paid one (config/thresholds.json#llm_pipeline.max_pages_in_llm_sample), not the whole site.

Then a second model joins those summaries together with the deterministic checks' findings and the site tree, and returns at most 10 recommendations. Each has a title, an explanation, a category and a priority from 1 to 5. Those ten become at most ten LLM-* findings.

Since 19 September 2026 that second model also receives a block of measured facts about your site: whether it has an llms.txt, a robots.txt and a sitemap, which types of structured data we saw and where, whether language variants are linked by hreflang, and on how many of the pages we looked at there is a canonical, a machine-readable author and an update date. Every item carries one of three states: present, absent, we did not check. Until then the model received only the list of problems we FOUND, and from their silence it inferred what the site lacked.

Finished recommendations then pass through a deterministic safeguard that watches for claims that the site lacks something we saw on it. Why it exists and what it does not cover is below, under When a recommendation says something untrue about your site.

From which follows the first surprise: these codes are not stable. The number at the end (LLM-STRUCTURE-01) is merely the position in that particular run. Run the audit again and the same content may carry a different number, or the recommendation may not come at all. So do not use an LLM-* code as a reference to a specific finding, unlike ACC-02, which always means the same thing.

Why they never get critical severity

An LLM-* finding cannot be critical. Never, not even when the model gave it the highest priority of 5. Priority 3 and above becomes a warning, lower becomes information, and that is where the ceiling sits.

This is not caution held in reserve. It is a response to something specific that happened on 1 September 2026, during an audit of our own site: the synthesis returned three recommendations with priority 4 and 5, and the report showed them as three critical findings alongside zero criticals from the deterministic checks. It looked as though the site were in a state of emergency.

But those three criticals were not three problems. Those three recommendations restated, in other words, warnings the rule engine had already raised, namely a missing meta description (STR-03) and missing structured data (SCH-01). The model described them more urgently and thereby promoted them, in the report, above the measurements they were based on.

We stress what the reason actually was, because it is easy to read wrongly: it is not distrust of the model's quality. The model made no mistake there; it described a real problem. It was about strength of evidence. A model reading content is weaker proof than a deterministic measurement, not stronger, and in this product critical severity is reserved for what we can measure: blocked crawlers, noindex, content behind a login, content that requires JavaScript.

The model's priority therefore still orders the recommendations among themselves, but it must not override measurement.

When a recommendation says something untrue about your site

It happens, and we are not going to hide it. In September 2026 a customer reported an audit in which three recommendations said untrue things about their site. One of them opened by stating that the site does not publish an llms.txt file, when the file was there and returned HTTP 200 at its address. The customer was right.

The deterministic checks had not got it wrong. The finding ACC-02 (no llms.txt) was not in that audit at all, precisely because the file existed. The untruth came from the model: it was given the list of problems we found, but nothing about what the site has, and from that silence it inferred absence.

We fixed it in three layers, and none of them is a guarantee on its own:

  1. The model is given the facts. The block described above, so that it need not guess about files and structured data. This is the main fix, because it addresses the cause.
  2. A safeguard behind the model. When a recommendation claims the absence of something we have on record as present, the model gets one attempt to rewrite it from the actual state. If that fails too, the whole recommendation is dropped. Nothing is written in on the model's behalf, and the other recommendations stay.
  3. A sentence on every LLM-* finding saying that what it states about your site may be inaccurate and asking you to check it. It holds even for an untruth that slips through both layers.

We measured the effect on two sites, three runs each: five and three untrue claims before the fix, zero after. That is a measurement on two sites, not a guarantee.

What the safeguard does not cover. Leaving this out would amount to marking our own homework:

  • It can only drop a recommendation for sites written in Czech or English. It runs in all nine languages our findings come in, but only for those two did a person write and measure the word list. In the other seven languages it still spots the violation and asks the model to rewrite the sentence, yet if even the last attempt fails, the recommendation goes through. That is deliberate: a wrongly dropped recommendation costs you good advice and nobody ever finds out, while an untrue one that slips through is still covered by the mandatory sentence on every recommendation. A language joins the two only after a native speaker reviews it, not after machine translation.
  • It matches words, not meaning. A recommendation that says the same thing the long way round passes.
  • It only drops against what we have on record. For anything we did not check it drops nothing. Since September 2026 the crawler fetches llms-full.txt alongside llms.txt, so for new audits we know whether your site has it too; for older audits that record is missing and the guard says nothing about the file.
  • On a sampled crawl it makes no claim about the whole site. A recommendation to “add structured data on the remaining pages too” therefore passes, even though we found some on the pages we looked at. That is deliberate: losing a true recommendation is worse, because you never learn that it was dropped.
  • It only concerns claims that something is missing. Other inaccuracies, such as a misread industry or a missed connection, get past it.

That is why every recommendation carries the sentence asking you to verify what it says about your site. It is not a disclaimer held in reserve; it is the one layer that holds where the first two fail.

How to tell a recommendation from the model

Until September 2026 one single thing set them apart from measured checks: the code prefix. Anyone who does not know the codes had no way of telling that one card is a measurement and the next one the model's judgement. Today it shows in three places:

  • A sentence at the end of the finding's text: “This recommendation was written by a language model reading the text of your pages. It is not a measured check, and we have no evidence for its effect. What it says about your site may be inaccurate, so please check it.” It lives in the finding's data, so you see it everywhere the finding's text appears: in the report, in the shared link, in the PDF and in the CSV export. E-mails carry no finding texts at all, only counts, so there is nothing to label there.
  • A “Language model” badge on the finding's card in the app, with an explanation for both the mouse and a screen reader. That one is in the language of your interface.
  • A sentence under the heading in the PDF, because in print there is nowhere to tuck an explanation into a tooltip.

The first of the three matters most: the badge lives only in the app, whereas the sentence is in the finding's data and travels with it wherever you take it.

How strong the evidence behind them is

For every other code in this manual we answer that question with a badge. Here a single badge will not do, because every recommendation is different. The model wrote it from what it read on your site.

What we can state is the upper bound: a recommendation from a model is never stronger evidence than the check it rests on. When a recommendation talks about structured data, everything under SCH-01 applies to it, namely that the effect on AI visibility is not documented. When it talks about links to sources, AUTH-01 applies, and there a measurement does exist.

A practical way to check for yourself: for each LLM-* finding, ask which deterministic check it relates to and read that check's page in this manual. The class of evidence is stated and sourced there. If a recommendation relates to no check at all, it is the model's own judgement over your text. It may be good, but we have no way to back it.

And one thing about content we can document, because a peer-reviewed experiment exists for it: Princeton (KDD 2024, 10,000 queries) measured that visibility in generated answers was lifted by quotations in the text (+41 %), concrete statistics (+33 %) and links to sources (+28 %). When a recommendation points that way, it has the strongest backing the field offers.

What the model cannot see, and what it may invent

Limits to know before rewriting anything on a recommendation's say-so:

  • It sees only the sample of pages. At most 10 or 30 depending on your plan. It knows nothing about the content of the remaining pages and may recommend something you already have elsewhere. Since September 2026 it also gets the block of measured facts alongside that sample, so at least about the site's files and structured data it need not guess.
  • It sees text, not appearance. It does not render the page, so it does not know what is visible on the first screen, what is hidden behind a click, or how large the page is.
  • It knows neither your field nor your customers. “Add a price list” may be wide of the mark for a company that never publishes prices on principle.
  • It is not deterministic. The same audit run twice over the same site may return different recommendations. That is a property, not a fault.
  • It can be wrong with confidence. The text sounds equally assured whether the claim is supported or not, unlike the rest of the report, where the badge tells you. Hence the safeguard and the sentence asking you to check: catching it completely is not possible.

That last point is why this page exists. Everywhere else in the report we tell you how strong the evidence behind a finding is. For LLM-* we cannot, and we consider admitting that more honest than attaching a badge it would not deserve.

What about text that tries to give the model orders

Worth mentioning, because it concerns the security of your report. Text from somebody else's site is untrusted input to us: anyone can write “ignore previous instructions” on a page and hope it reaches the output.

We defend against that in three layers: the page content is delimited in the prompt, the model is explicitly told never to execute instructions inside those delimiters, and its output is validated against a strict schema. When no valid structure comes back, it is retried once and then the page is marked as unprocessed; invalid output is never used.

In practice: even if text on somebody's page confused the model, the worst that happens is an oddly worded recommendation in the report. Nothing is executed and nothing is written anywhere.

What to do with the recommendations

An order that makes sense and matches the strength of evidence:

  1. Close the necessary conditions first. ACC-01, ACC-03, ACC-05, ACC-07: while those are open, the rest is moot.
  2. Then the deterministic findings, where the class of evidence is stated and sourced.
  3. Read LLM-* recommendations as editorial notes. They are most useful where they tell you something about content that no measurement catches: that the text does not answer the question its heading promises, or that a concrete figure is missing.
  4. Check what it says about your site before you change anything on its say-so. If it claims you do not have something, look at whether it is really missing. That is the one extra thing we ask of you with these findings, and it is exactly where the model can get it wrong.
  5. Discard whatever does not match reality. The model does not know your field. A recommendation you know to be wide of the mark is wide of the mark, and nothing is spoiled by dropping it.

And a note on language, since it comes up often: on language, LLM-* are the one exception to the rule that governs the rest of the report. Findings from deterministic checks are translated, so they show in the language of your interface whatever the audited page was. A recommendation from the model we cannot translate (no template exists for it, and translating would mean another model call), so it stays in the language of the audited page and carries a badge with that language's code. Whoever audits a German site from an English interface gets German recommendations and can hand them straight to a German content editor.

Sources