People treat sitemaps like a ranking cheat code. They’re not. They’re a hint list: “these URLs exist; please consider them.”
Google’s sitemap overview says the quiet part:
If your site’s pages are properly linked, Google can usually discover most of your site.
And then — usefully — when you might need one: large sites, new sites with few external links, lots of media/news you care about in Search. Also when you might not: roughly ≤500 important pages, strong internal linking, no special media needs.
If you only remember one rule from me: a smaller honest sitemap beats a bloated dishonest one.
urlset vs index
Per the sitemaps.org protocol, you’ve got:
- urlset — list of URLs
- sitemap index — list of other sitemap files (size splits, sections, locales)
Both fine. HTML at /sitemap.xml pretending to be a map is not fine. That soft-404 costume burns hours.
Size ceilings in the protocol (50,000 URLs / 50MB uncompressed per file) are why indexes exist. Don’t invent your own “unlimited” file and hope.
Discovery paths
- robots.txt:
Sitemap: https://example.com/sitemap.xml - Search Console / Bing submit
- Conventional
/sitemap.xml(helpful, not sufficient alone)
Google’s build docs also cover submitting and formats — start from Build and submit a sitemap once the file is real XML.
Check both file + robots declaration: Sitemap Checker, Robots.txt Checker.
Include / exclude (the part teams skip)
Include: preferred host, HTTPS, 200s, indexable, actually important.
Exclude: thank-you, cart, account, internal search, facet explosions, staging hosts, URLs that only redirect (list the destination), noindex templates, anything robots.txt disallows.
If the sitemap and noindex disagree, you trained Search Console to distrust you. Canonical story should agree too — Google treats sitemap inclusion as a weak canonical signal in their canonicalization guide, and they warn against conflicting canonical techniques.
lastmod, priority, changefreq
Be honest or omit.
priority and changefreq have been shrugged at for years in practice; don’t spend engineering time generating decorative fiction.
<lastmod> helps when it’s true. Fake “updated daily” trains systems to ignore your dates. Google’s sitemap docs discuss lastmod in the build guidance — the spirit is accuracy, not theater.
Validation after deploys that touch URLs
- Fetch
/sitemap.xml+ every robotsSitemap:URL - Confirm XML, not HTML
- Sample locs for wrong host / HTTP leftovers / UTM junk
- Spot-check five URLs for status, canonical, noindex
- Then submit or re-submit in Search Console
Steps 1–3: Sitemap Checker. Step 4: Noindex, Canonical.
Search Console messages, decoded
Couldn’t fetch — wrong URL, auth, non-XML, edge timeout. Fix fetch first.
Sitemap is HTML — theme soft-404. Fix the route.
URL not allowed — often host mismatch or robots disallow.
Submitted URL noindex / wrong canonical — stop resubmitting the contradiction; fix the page or remove the loc.
Discovered but not indexed — discovery worked. Quality/demand problem now. Submitting harder won’t save thin pages.
Migrations
New host → new sitemap with new locs only. Redirect old URLs. Don’t keep advertising a dead map on the old host forever. Change-of-address in Search Console where it applies; sitemaps alone aren’t a migration plan.
Myths
More URLs ≠ more traffic. Sitemap ≠ guaranteed indexing. Image/video sitemap extensions help those surfaces (Google’s overview covers why media/news sites care) — they don’t fix empty HTML. CMS default sitemaps are often full of junk after a year of plugins; always sample.
For the whole technical picture: audit. For sitemap-only: Sitemap Checker. Spec roots: sitemaps.org + Search Central sitemaps.