SEO-34
Declared page language does not match the content
What the check measures
The finding is raised in one deliberately narrow case: the page declares <html lang> with a valid code other than cs, but its text contains Czech diacritics. The report gives both the declared and the guessed language.
Why only Czech, when we can write findings in nine languages? Because guessing a language from text is only reliable for Czech: the diacritics are an unambiguous signature. For other languages, widening the comparison would mainly create chances to guess wrong, and a finding that occasionally accuses a correctly configured page is worse than one that stays quiet.
We do not measure the language, though; we estimate it by statistical analysis of the text. We count what share of all letters on the page are the characters č ď ě ň ř š ť ů ž: the ones that, apart from Czech and Slovak, none of the nine languages we can write findings in uses. If there are at least three of them and they make up 2.5 % or more of all letters, we call the text Czech. The threshold and the character list live in apps/.
The threshold comes from measurement, not from a hunch. Genuinely Czech pages have a density of 4.43–4.57 %; foreign-
Statistics can be wrong and we say so out loud. A false finding appears where the Czech characters are out of proportion to the length of the text: a short foreign-
That is why the optional step “Language check by AI model” exists. Turn it on and a language model judges the language of the doubtful pages: it reads the text instead of counting characters. We only ask it where the answer can change something: on pages where the finding would be raised, on pages whose density falls in the 1–4 % band (where the two measured groups overlap), and on a small sample of pages with no Czech signal at all.
The result lands before the finding is created, not as a later correction: if the model confirms the mismatch, the finding is raised; if it refutes the estimate, no finding appears at all. If the model fails, statistics apply and the finding behaves as if the step never ran. And only with verification can this check find a mismatch outside Czech: a German page declared as French, say, which statistics cannot see.
What the check does not do:
- It will not spot a German page declared as French, or any other pair without Czech.
- It does not run on a page with no
lang: without a valid code there is nothing to compare. That isSEO-14. - It will not spot the inverse: a Czech declaration on English text.
- It cannot see nested language passages. A Czech quotation inside an English article can trigger it even though the page is fine.
The finding is page-level; severity is warning.
How strong the evidence is
We recommend it because it does no harm or has some other benefit, but we promise nothing about whether it makes language models cite you. Nobody has demonstrated that yet.
For visibility in AI answers we have no documented effect. And for search it is weaker than expected; Google has it in its own documentation:
Google uses the visible content of your page to determine its language. We don't use any code-level language information such as lang attributes, or the URL.Google Search Central, multilingual sites
So a wrong declaration probably will not confuse Google: it reads the language from the text.
What it will reliably confuse is a screen reader. It takes lang literally and chooses a voice model from it. A Czech sentence read with English pronunciation is unintelligible to a blind user. That is not an inconvenience, it is a barrier. It is a WCAG criterion and it is the hardest reason this finding has.
The second impact is on the browser: lang drives hyphenation and the translation offer. A Czech page marked as English will offer visitors a translation into Czech, from Czech.
The practical conclusion: fix it, it is one attribute. But the reason is with people, not with search engines, and we will not tell you otherwise.
How to fix it
First work out which case you have. The finding has two entirely different causes and the second is the more serious:
- A misconfigured template. The site is Czech but the template carries
lang="en"from its original. A one-line fix that clears the whole site. - A multilingual site with a shared template. The English version correctly carries
lang="en", but part of the content on it stayed Czech: an untranslated article, an untranslated footer, untranslated product names. Here the defect is in the content, not in the attribute, and rewritinglanghides it instead of fixing it.
<html lang="cs">- Czech is
cs, notcz. - Every language version carries its own
lang: generate it from the same source that decides the content language, not hard-coded in the template. - Mark a foreign-
language passage inside the page separately with langon its wrapper. A screen reader switches pronunciation and the finding clears correctly, not by a workaround.
<p>Google puts it this way: <span lang="en">There are no additional requirements.</span></p>What the report says about it
Finding description
`<html lang>` declares "[declared]", but the page's visible text contains distinctive Czech diacritics, a strong signal that the content is actually "[guessed]". An incorrect `lang` is a confusing signal for search engines and assistive technology alike (screen readers pick the wrong pronunciation). This is a standard recommendation for classic search; we have no evidence it affects visibility in AI assistant answers. We do not measure the content language, though. We estimate it by statistical analysis of the text until the optional audit step “Language check by AI model” confirms it. The estimate can be wrong, typically on a foreign-
Recommendation
Fix `<html lang="…">` to match the page's actual content language, or (if the page is genuinely multilingual) split the content into separate URLs with their own `lang` and reciprocal `hreflang` (SEO-05).
Sources
- Google Search Central: Managing multi-regional and multilingual sites (accessed 2026-09-12)
- W3C: Language tags in HTML and XML (accessed 2026-09-12)
Text verified 2026-09-13