ACC-06

Site sitemap (sitemap.xml)

WarningCrawler access

What the check measures

We look for the sitemap at https://your-domain/sitemap.xml, in the sitemap index that path may point to, and at the addresses declared in robots.txt - always on the audited domain only. Over whatever is found, we judge five things in this order: whether the file exists, whether it is a readable sitemap (valid XML with a <urlset>/<sitemapindex> root and at least one <loc>), whether the sitemap lists addresses we could not reach or that return an error, whether it lists addresses your own robots.txt forbids us to crawl, and whether there are live pages missing from the sitemap.

The finding always tells you which of those five happened, right in the title. It used to carry one sentence covering them all (“missing or invalid sitemap.xml”), which on a site with a working sitemap was simply untrue; fixed on 19 September 2026.

The first two conditions always apply. The other three are evaluated only on a full crawl, never on a sample, a deliberate decision worth explaining. Over a sample of a few dozen pages, a mismatch would come out almost every time, not because of your site but because of us: we simply never fetched that URL. A finding that cannot tell a defect in your site from a limit of our own crawl is one we would rather not show.

What the check does not do:

  • It does not validate against the sitemaps.org schema. Only the rough shape is verified: a <urlset>/<sitemapindex> and at least one <loc>. A malformed <lastmod> or an invalid priority passes.
  • It does not read a sitemap index without limit. For a large index we read only part of the child files (the cap lives in config/thresholds.json); when that happens, the “pages missing from the sitemap” branch stays silent, because it would be a statement about us rather than about your site.
  • It does not compare <lastmod> against actual page changes. A sitemap claiming every URL changed today passes, even though that is precisely the signal that wastes crawl budget.
  • URLs that are merely the source of a redirect do not count as “missing from the sitemap”. An old URL does not belong there; its destination carries the entry. And the other way round: an address from the sitemap that redirected us to a page the audit did crawl does not count as “we could not find it” - a redirect is not a fault.
  • Pagination (/section/page/2/) and variants of the same address with a query parameter do not count as missing from the sitemap. A sitemap holds the canonical address, not every variant of it; advising you to add those would make your sitemap worse.
  • It does not verify that the search engine actually took the URLs from the sitemap. A sitemap is an invitation, not a guarantee, and Torumata cannot see into the index of Google or Seznam. How many pages are really in the index is answered by Google Search Console and, for Czech sites, by Seznam Webmaster.

The finding is site-level; severity is warning.

How strong the evidence is

Effect not demonstrated

We recommend it because it does no harm or has some other benefit, but we promise nothing about whether it makes language models cite you. Nobody has demonstrated that yet.

Two things get routinely merged into one in GEO articles, and they need separating here.

For discovering pages, sitemaps work, and Google documents it: it is how you tell a search engine which URLs exist on your site and when they last changed. They help most on large sites, on new sites with few inbound links, and for content that is poorly linked internally.

For visibility in AI answers we have no documented effect, which is why this check sits in the “not demonstrated” class. No operator of a language model declares that it reads sitemaps, and no measurement shows that a site with a sitemap is cited more often than a comparable site without one.

Two things worth stating outright, because they get claimed about sitemaps and do not hold:

  • A sitemap does not guarantee indexing. Google says so literally: it does not guarantee that all the items in your sitemap will be crawled and indexed. It is a hint, not a command.
  • A sitemap is not a ranking factor, as far as we know. Google does not list it among its ranking systems; we do not claim the opposite either, since we could not document that. The documented role of a sitemap is discovering URLs, not ordering them.

So why do we recommend it? Because it costs almost nothing and helps where it genuinely matters: getting the page found at all. Being indexable is a necessary condition for appearing in AI (ACC-04, ACC-07), and being discovered is the step before that condition can even apply. That is an honest justification. Promising more would mean claiming something we cannot demonstrate.

In practice: the part of the finding about a mismatch is worth more than the part about the file existing. A sitemap listing hundreds of URLs that return 404 is not merely useless; it is a bad hint that you are sending the crawler yourself.

How to fix it

Do not write the sitemap by hand. A hand-maintained file goes stale before anyone notices, and that is exactly how the worse half of this finding comes about: a sitemap claiming things that are not on the site. Let it be generated automatically when content is published, from the same source as your navigation.

The minimum valid file looks like this:

<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
  <url>
    <loc>https://your-domain/</loc>
    <lastmod>2026-09-12</lastmod>
  </url>
  <url>
    <loc>https://your-domain/article</loc>
    <lastmod>2026-08-30</lastmod>
  </url>
</urlset>

Then point at it from robots.txt on its own line, which belongs to no User-agent group:

Sitemap: https://your-domain/sitemap.xml

What belongs in a sitemap and what does not:

  • Canonical URLs only. Parameterised versions, print variants and duplicates have no place there: you would be sending the crawler at pages you yourself label as non-preferred.
  • Only URLs returning 200. Every 404 and every redirect in a sitemap is a hint that does not hold.
  • No pages carrying noindex. The sitemap says “index me”, the meta tag says the opposite. Two contradictory claims about one URL.
  • <lastmod> only when you know it. The date of the last real content change, not the date the file was built. When everything says today, it stops being information.
  • Above 50,000 URLs or 50 MB, split the file and add a sitemap index. Those limits come from the sitemaps.org protocol, not from us.

When the finding reports a mismatch, the title tells you which direction. Addresses the sitemap lists but we could not reach are usually old paths left after a site rebuild, or pages nothing links to internally; the latter is also SIT-01. Live pages missing from the sitemap tend to be content generated outside the main template: tags and archives. For each such address, decide first whether it really belongs in the sitemap - pagination and filtered listings do not, which is why we no longer report them to you.

What the report says about it

Finding description

We found no usable sitemap, or it doesn't match reality. We look for it at /sitemap.xml (including a sitemap index pointing to partial files) and at the addresses listed in robots.txt, always on the audited domain. Either nothing loaded, or the file couldn't be read as a sitemap, or it doesn't match the pages actually found by the crawl (it lists URLs this audit did not visit, or is missing URLs the crawler did discover). “We did not visit it” is not the same as “it is not on the site”: a single audit need not cover every address.

Recommendation

Generate a valid sitemap.xml listing all publicly accessible pages (canonical URLs, `<lastmod>` where known) and reference it from robots.txt (`Sitemap: https://yourdomain/sitemap.xml`). Splitting it into a sitemap index with partial files is fine, and so is keeping it at a different address as long as robots.txt points to it. Keep it auto-generated on publish, not maintained by hand. A sitemap is an invitation, not a guarantee: Torumata does not check whether the search engine really took the addresses from it. The real state is shown by a `site:` query in the search engine, by URL Inspection in Google Search Console and, for Czech sites, by Seznam Webmaster.

Sources

Text verified 2026-09-19