SCOPE-01
The sitemap lists a different number of addresses than the audit found by following links
What the check measures
The size of the site is estimated from several independent sources: how many URLs the sitemap lists, and how many distinct addresses we discovered by following links. The finding is raised when those two measured numbers diverge by more than one and a half times (the larger over the smaller is 1.5 or more). The report gives both.
The 1.5 threshold lives in config/ under consistency. and it is deliberately tolerant. A small discrepancy between a sitemap and reality is normal; the finding should fire only where the gap is large.
What the check does not do:
- It does not say which source is right. The sitemap may be inflated with old URLs, or link-following may have hit a limit. Both look the same.
- It does not list the specific URLs that differ, only counts. Specific mismatches are reported by
ACC-06andSEO-07. - It works from measurements, never from an estimate. Our estimate of the whole site can be higher than the number of addresses we discovered (we add one conservative extrapolation step) - that computed number enters neither this finding nor the cross-check between sources. Until 19 September 2026 it did, and the finding then claimed we had found more addresses than we had actually seen.
- It does not count addresses that lead to a file. Images, video, audio, documents, archives, stylesheets, scripts and fonts we recognise from the extension in the address or from the content type the server reported, and we leave them out of the discovered pages. Otherwise a gallery where clicking a thumbnail opens the photo itself would inflate the count by hundreds of addresses; on one audited site that was two thirds of all the “discovered pages”. The list of extensions lives in
config/and is deliberately closed: it holds only formats where the answer is unambiguous, sofile- url- extensions. json .phpand.htmlare not among them. - It does not conclude that your sitemap is stale. The difference may simply be pagination and parameterised addresses we discovered, which do not belong in a sitemap.
- It does not run when there is only one source of the estimate (a site with no sitemap).
The finding is site-level; severity is warning.
How strong the evidence is
We recommend it because it does no harm or has some other benefit, but we promise nothing about whether it makes language models cite you. Nobody has demonstrated that yet.
For visibility in AI answers there is no documented effect, and this finding is not really about your site in the usual sense. It is a finding about the reliability of the picture we formed of your site.
Why we tell you rather than keep it to ourselves: because it has consequences for the rest of the report, and it is more honest to state them than to hide them.
When the estimates diverge, we do not know for certain how large your site is. The audit walks a sample, and some checks switch themselves off when they cannot claim anything about the site as a whole: SIT-01 (orphaned pages), SEO-12 (unlinked pages) and part of ACC-06 stay quiet when the crawl was not complete. A large spread in the estimate therefore means you may have received fewer findings than a complete pass would have produced.
The second consequence is commercial and concerns money: the scope of the crawl feeds into the price of the audit. When a sitemap lists ten times more pages than link-following found, the estimate is easy to get wrong.
The practical conclusion: this is not a defect of your site that you must fix. It is the information “do not fully trust this number”, and a pointer to where to look when the scope of the report does not match what you know about your site.
How to fix it
First find out which of the two sources is off. Two questions decide it:
| What you see | What it probably means |
|---|---|
| The sitemap lists far more than we found | The sitemap is inflated with old or non-existent URLs, or nothing links to most of the content and the only way in is through the sitemap. |
| Link-following found far more than the sitemap lists | The sitemap is stale, is generated from only part of the site (typically articles only, without products and categories) - or it is fine and we discovered addresses that do not belong in it: pagination, filtered listings and parameterised variants. Before you start adding, look at what those extra addresses actually are. |
Then, accordingly:
- An inflated sitemap → regenerate it from existing, indexable URLs only. The specific URLs we could not find are listed by
SEO-07. - An incomplete sitemap → check what generates it. It is usually a plugin that knows only one content type.
- Content reachable only from the sitemap → add links to it from the site. That is also
SIT-01andSEO-12, and it matters more than the mismatch of numbers. - A sitemap not generated automatically → that is the root of most of these cases. A hand-made sitemap goes stale before anyone notices.
And one thing you may run into on the way that is worth more than this finding: when link-following finds markedly fewer pages than you actually have, the problem is navigation, not the sitemap. Content with no route from the site itself is harder for people and search engines alike to find; the sitemap merely masks it.
What the report says about it
Finding description
The sitemap lists [sitemapPages] addresses; by following links we discovered [linksPages]. Both numbers are measured; neither is an estimate. Addresses that lead to a file rather than to a page are left out of the discovered count: images, video, documents, archives and similar formats we recognise from the extension in the address or from the content type the server reported. Otherwise a gallery where clicking a thumbnail opens the photo itself would inflate the count by hundreds of addresses. The difference alone does not mean your sitemap is wrong: pagination, variants with query parameters and `noindex` pages do not belong in a sitemap, so a crawler routinely discovers more addresses than a sitemap lists. In the other direction the sitemap may be more complete than our own list, because an audit has a limited scope. It is still worth finding out what the difference is made of. The site-size estimate also drives the pre-filled price of your next audit, so the further apart the sources are, the less reliable that estimate is.
Recommendation
Compare the two sets and decide, for each extra address, whether it belongs in the sitemap. Pages with their own content that you want indexed do, in canonical form. Pagination, filtered and parameterised variants and `noindex` pages do not - do not add those, you would make the sitemap worse. Remove addresses that are in the sitemap but no longer on the site. The safest route is to have the sitemap generated automatically when content is published and to point to it from robots.txt (`Sitemap: https://
Sources
- Google Search Central: Learn about sitemaps (accessed 2026-09-12)
Text verified 2026-09-22