seotools.pro Blog

How 11 of 19 Sitemaps Break the Rules

We audited 21 live sitemaps against the protocol. 8 were clean and the worst offenders sell SEO software. 7 lessons on the file nobody rereads.

We audited 21 live sitemaps against the protocol, and the worst offenders sell SEO software.

Only 8 came back clean. Ahrefs ships no lastmod at all, and Semrush ships none either.

Your sitemap is the one file on your site nobody rereads. You generate it once, it validates as XML forever, and no tool ever tells you it has quietly stopped being useful.

So we re-fetched 21 of them on 6 September 2026 and read what was actually inside. Here are the 7 lessons, and every one of them applies to your file too.

Assume yours is broken until you look

We requested 21 sitemaps and 19 answered. Etsy returned a 404 at the URL it declares in its own robots.txt, and GitHub refused our request outright.

Your own file could be doing either of those right now and nothing would tell you.

Of those 19, 8 had nothing wrong with them at all: Booking.com, Canva, eBay, Figma, Investopedia, Moz, Vercel and ours.

The other 11 each carry something that is ignored, out of scope, or missing. None of it is a syntax error.

Every one of these files is valid XML and would pass a validator without a word of complaint. That is exactly why the problems survive, and why yours has survived too.

So assume your own file is broken until you have opened it and read it yourself. Nothing in your stack is going to tell you otherwise.

Your build generates it, your validator passes it, and your rank tracker never mentions it. If you do not look, nobody looks.

Cut the fields nobody reads

The most common finding, and the most pointless.

<changefreq> and <priority> are advisory fields in the sitemaps.org spec, and Google's own documentation states it ignores both. 5 of the 19 still send changefreq and 4 still send priority.

Cloudflare sends 896 of each. Shopify sends 557 of each. The Guardian sends 424 changefreq values. Zapier sends 36 of each. HubSpot sends exactly 1 of each, which is somehow the strangest result in the survey.

Cloudflare's file carries 1,800 elements that no major search engine reads. They cost nothing but bytes.

The point is not the bytes. It is that the file was generated years ago by a tool with those defaults, and nobody has opened it since. Your generator had defaults too, and you probably did not choose them either.

Give every URL a real lastmod

The field that does get read is the one most often missing.

<lastmod> tells a crawler which of your pages changed, so it can spend its budget on those. 6 of the 19 have none: Ahrefs, Semrush, Stripe, Notion, Healthline and Zapier.

Nothing breaks without it. A crawler simply gets no signal about what is new and falls back on its own guess, which on a large site means your updated pages queue behind thousands of unchanged ones.

Note the shape of that list: a payments API, a news publisher, an automation platform and two SEO vendors. Size and sophistication predict nothing here.

If you want to see what your own file declares before you fix it, the XML Sitemap Generator will walk the indexes and show you what a crawler receives.

That is free. A sitemap problem shows up as pages quietly dropping out of the index, so keep a ranking history beside it — Semrush and SE Ranking both keep one, and the drop is visible there weeks before anyone notices.

Discount advice the adviser has not taken

Of the 3 largest SEO software companies, 1 has a clean sitemap.

Moz is clean. Ahrefs ships no lastmod. Semrush ships no lastmod and 7 URLs pointing at a different host than the sitemap serving them — entries a crawler is entitled to ignore, because a sitemap may only speak for URLs on the host it lives on.

We are including ourselves in this comparison, so: unmiss.com's index has 7 entries and is clean. That is a small file and not much of a boast.

The relevant part is that somebody checked it. That is the entire difference between the two halves of this table.

You can put yourself in the top half this afternoon. Nothing on this list needs a project, a budget or a vendor. It needs someone to open the file.

Work with us
Most sitemaps are quietly broken

11 of the 19 sitemaps we checked break the rules, including several from companies that sell SEO. Nobody rereads these files. We do, along with everything else no one has looked at since launch.

  • A full technical audit, read by a human
  • Developers who fix what the audit finds
  • SEO and GEO once the basics are clean
  • First consultation is free
Book a free consultation → See our SEO and GEO services →

Open the file, do not trust the status

The single strangest result. wikihow.com/sitemap_index.xml returns a 200, valid XML, correct root element — and 0 URLs.

An empty sitemap index is indistinguishable from a healthy one in every automated check that stops at "did it parse". It returns 200. It validates. It contains nothing.

wikiHow is a site of hundreds of thousands of articles that ranks extremely well, so this is not costing them their business. It is a good illustration that a sitemap can be entirely absent in effect while appearing entirely present, and that nobody would find out unless they counted the entries.

So count yours. It is one line of shell, and it separates "we have a sitemap" from "our sitemap has our pages in it".

Those two sentences sound the same in a status meeting. Only one of them is true on most sites, and you will not know which until you look.

Scope it to its own folder

The protocol restricts a sitemap to URLs at or below its own path. A sitemap at /sitemaps/news.xml may list /sitemaps/… and nothing above it.

The Guardian's news sitemap lists 424 URLs outside its own directory — ordinary article URLs at the site root, submitted from a file that has no authority over them.

In practice Google is forgiving about this when the sitemap is declared in robots.txt, which the Guardian's is. But it is the kind of thing that works until a crawler decides to apply the rule, and there is no upside to being out of spec.

So put your sitemap at the root, or move it up to cover everything it lists. It costs you one redirect and removes an entire category of silent failure.

Check it after every deploy

The instruction, and the reason this survey exists.

Everything above is invisible to the person who owns it. There is no error, no ranking drop you could attribute, no alert. The file was right when it was generated and drifted afterwards.

So put it in the deploy. Fetch the sitemap, count the entries, assert the count is what you expect, and check the fields you care about are present. We do this in our own build now — an earlier version of this article had to report that we did not, and that our own file failed a check we were running on other people's.

Then look at what the crawler does with your file. Our Website Audit flags sitemap URLs that carry a noindex tag, which is the most expensive version of this problem — asking to be indexed and declining in the same breath. Canva has that one live right now.

Free, no account
See what your sitemap actually declares

Every file in this survey is valid XML. The problems are in what the file says, not whether it parses — which is why 11 of 19 have been wrong for years without anyone noticing.

  • Walks your sitemap index and counts what a crawler receives
  • Pairs with the audit's noindex-in-sitemap check
  • Free, no signup, no card
Check your sitemap free

What we could not measure

And here is what this audit does not show. A couple of the tools named on this site are partners of ours — if you buy through those links we earn a commission, and you do not pay a cent more.

We did not measure any effect on rankings. Nothing here connects a missing lastmod to a position, a crawl rate or a visit. These are conformance findings — the file does not say what the protocol says it should — and conformance is not the same as consequence.

That Google ignores changefreq and priority is Google's own published position, not something we tested. We counted the fields; we cannot observe what any crawler does with them.

We audited the top-level sitemap or index only, not every child file. A clean index can point at a broken child, and 8 of our 19 are indexes, so "clean" here means the entry point is clean. Booking.com's index alone expands to hundreds of thousands of files, which we did not walk for this survey.

Two sitemaps could not be read at all. Etsy returned a 404 at the URL its own robots.txt declares, which is itself a finding we have not chased.

GitHub returned a 406 to our user agent, which is a block on us rather than a fault in their file.

The sample is 21 sites we chose, not a random sample, and it is weighted towards companies an SEO audience would recognise. The method is scripts/survey-sitemaps.mjs in our repo and the raw output is in evidence/surveys/.

Frequently asked questions

How many sitemaps had problems?

11 of the 19 we could read, measured 6 September 2026. The clean 8 were Booking.com, Canva, eBay, Figma, Investopedia, Moz, Vercel and seotools.pro. Every file in the survey is valid XML — the problems are in what they declare, not whether they parse.

Does Google use changefreq and priority?

No. Google's own documentation states it ignores both. 5 of the 19 sitemaps still send changefreq and 4 still send priority — Cloudflare emits 896 of each, Shopify 557 of each, and HubSpot exactly 1 of each.

Does lastmod matter?

It is the field a crawler does read, and it tells the crawler which pages changed so it can prioritise them. 6 of the 19 have none at all: Ahrefs, Semrush, Stripe, Notion, Healthline and Zapier. We did not measure any ranking effect from its absence.

Can a sitemap list URLs from another folder?

The protocol scopes a sitemap to URLs at or below its own path. The Guardian's news sitemap lists 424 URLs above its directory, and Semrush's lists 7 on a different host entirely. Google is forgiving in practice when the file is declared in robots.txt, but there is no upside to being out of spec.

Can a sitemap be empty and still look fine?

Yes, and one in this survey is. wikihow.com's sitemap index returns a 200, parses as valid XML, has the right root element and contains 0 URLs. Every automated check that stops at "did it parse" passes it.

What are the sitemap size limits?

50K URLs and 50 MB uncompressed per file. None of the top-level files in this survey came close, though IKEA's product sitemaps — measured separately — reach 47.6 MB each because of the hreflang blocks inside them.

Do SEO tool vendors have clean sitemaps?

One of the three largest does. Moz is clean; Ahrefs ships no lastmod; Semrush ships no lastmod plus 7 URLs on another host. Ours is clean and has 7 entries, which is a small file rather than an achievement — the difference is that it gets checked.

How do I check my own sitemap?

Fetch it, count the <loc> entries, and assert the count is what you expect — then check that lastmod is present and that every URL is on the sitemap's own host and path. Put it in your deploy, because nothing else will tell you when it drifts.

The one number to take away

0.

That is how many URLs are in wikiHow's sitemap index, on a site with hundreds of thousands of articles — and how many automated checks would tell them, because the file returns 200 and parses perfectly.

That is the shape of every finding here. Nothing in this survey is broken in a way anything reports. Cloudflare's 1,800 ignored fields, Ahrefs' missing lastmod, the Guardian's out-of-scope URLs, Semrush's off-host entries — all valid XML, all silent, all years old.

The sitemap is the last file anyone opens and the first one to stop being true. Ours was wrong until we wrote a check for it, and the check took an afternoon.

Go and count the entries in yours. If the number surprises you, that is the finding.

Measured 6 September 2026. We fetched 21 top-level sitemaps — discovering the URL from robots.txt where a site does not use the conventional path — and checked entry counts, lastmod presence and format, changefreq and priority usage, host and path scope, and the 50K-URL and 50 MB protocol limits. 19 answered. The method is scripts/survey-sitemaps.mjs in our repo and the raw results are in evidence/surveys/, so every figure here is re-runnable.

Blog