SEO-14

Lang attribute missing or invalid

WarningStructure

What the check measures

On every live page we read the lang attribute on <html> and verify its shape: whether it looks like a valid BCP 47 language tag. The finding is raised when the attribute is missing, empty, or obviously malformed (cz, czech, en_US with an underscore).

Only the shape is checked, not the truth of it. A Czech page declaring lang="de" passes this check: it has a valid code. A mismatch between declaration and content is reported separately by SEO-34, which can reliably recognise only Czech.

What the check does not do: it does not check lang on nested elements (a foreign-language quotation inside an article), it does not compare against the content language, and it cannot see an attribute added by JavaScript. The finding is page-level; severity is warning.

How strong the evidence is

Effect not demonstrated

We recommend it because it does no harm or has some other benefit, but we promise nothing about whether it makes language models cite you. Nobody has demonstrated that yet.

For visibility in AI answers we have no documented effect, hence “effect not demonstrated”. As with STR-05, though, it is worth separating who the attribute actually serves, because the answer is not where people expect.

The main beneficiary is a screen reader. It picks pronunciation from lang. A Czech sentence read by an English voice model is unintelligible to a screen reader user: not inconvenient, unintelligible. It is a WCAG criterion and it is the hardest reason this finding has.

The second is the browser: lang drives hyphenation, spell-checking and the translation offer.

For search engines it is weaker than claimed, and Google says so unusually bluntly. Its documentation on multilingual sites states:

Google uses the visible content of your page to determine its language. We don't use any code-level language information such as lang attributes, or the URL.Google Search Central, multilingual sites

So our finding reports something Google explicitly says it does not use. We recommend it for people and browsers, not for your position in results. And if anyone tells you lang improves your SEO, you now know where to check.

How to fix it

One attribute on the document's opening element:

<html lang="en">
  • Czech is cs, not cz. cz is a country code, not a language, and it is the commonest mistake on Czech sites.
  • A hyphen, not an underscore, when you add a region: en-GB, not en_US.
  • Do not add a region you do not need. en is right for English text; en-GB is not wrong, but it adds nothing.
  • A multilingual site carries its own lang on each language version: one hard-coded lang="en" in a template shared with another language is worse than none, because it states something untrue.

When a longer passage on the page is in another language, mark it separately; a screen reader switches pronunciation mid-text:

<p>The Czech original reads: <span lang="cs">Žádné další požadavky neexistují.</span></p>

What the report says about it

Finding description

`<html lang>` has the value "[value]", which does not look like a valid language code (BCP 47). Search engines and assistive technology read it directly, and a wrong or missing attribute is a confusing signal. This is a standard recommendation for classic search engines (Google/Bing); we have no evidence it affects visibility in AI assistant answers.

Recommendation

Set `<html lang="…">` to a valid BCP 47 code matching the actual language of the page content (e.g. "cs", "en", "de-AT").

Sources

Text verified 2026-09-12