Free XML Sitemap Checker & Validator
We find your sitemap, parse it, validate it against Google’s rules, and then actually fetch a sample of the URLs inside it to confirm they resolve. Most sitemap checkers stop at the XML.
- Validated against Google’s rules
- Live URL sample
- Index files supported
Why check your XML sitemap?
Faster discovery
A clean sitemap is how new and deep pages get found. Without one, discovery depends entirely on internal links.
Live URL testing
We do not stop at parsing XML. We fetch a spread sample of your URLs and report 404s, redirects and noindex tags for real.
Google’s actual rules
The 50,000 URL and 50MB limits, cross-domain URLs, duplicates, malformed lastmod dates and http:// entries that redirect on every crawl.
Index file breakdown
We follow sitemap indexes and report each child separately, so you know which section of the site has the problem.
How it works
Enter your domain
Or paste a sitemap URL directly if you already know it.
We find your sitemap
We check robots.txt first, then nine standard sitemap paths, and follow any index file we find.
Validation and live sampling
The XML is validated against Google’s rules, then a spread sample of URLs is fetched to confirm they actually resolve.
Review the report
A score, errors and warnings with examples, per-child breakdown, and a downloadable URL list.
Every finding is measured live. Fix, re-run, and watch it clear — no signup, no limits.
Why this matters
A sitemap is how you tell search engines which pages exist and when they last changed. Without one, discovery depends entirely on crawlers following links — which means new pages take longer to be found, and pages nothing links to may never be found at all.
The failure mode people miss is subtler than a missing file. A sitemap full of redirects, 404s and noindexed URLs teaches Google that the file is unreliable, and it gets crawled less carefully as a result. That is why this checker does not stop at parsing the XML; it fetches a spread of the URLs inside it and reports what actually came back.
What gets validated
- Discovery — is the sitemap declared in robots.txt, or did we have to guess a standard path?
- XML validity — correct root element, sitemaps.org namespace, well-formed structure.
- Google’s hard limits — 50,000 URLs and 50MB uncompressed per file.
- Duplicates — the same URL listed more than once wastes crawl budget.
- Cross-domain URLs — ignored by Google unless the sitemap is cross-submitted in Search Console.
- Protocol consistency — http:// URLs that will redirect to https:// on every crawl.
- lastmod accuracy — valid W3C datetime format, and how much of the file actually has dates.
- Live URL status — a spread sample fetched for real, checking for 404s, redirects and noindex directives.
Why the live sample matters
Parsing XML tells you the file is syntactically fine. It tells you nothing about whether the URLs inside it are any good.
The three problems we find most often are all invisible to a pure XML validator. Redirect chains, where the sitemap lists old URLs that 301 to new ones — every crawl wastes a request. Broken URLs left behind after a page was deleted but the sitemap plugin never regenerated. And noindexed pages listed in the sitemap, which sends Google directly contradictory instructions: “crawl this important page” and “do not index this page”.
We sample rather than test everything, because fetching 50,000 URLs would take hours. The sample is spread evenly across the file rather than taken from the top, which is how systematic problems surface.
Sitemap index files
Once a site passes a few thousand URLs, the normal pattern is a sitemap index — one file that points at several child sitemaps, typically split by content type. WordPress plugins do this automatically.
The checker follows the index, fetches the child sitemaps, and reports on each separately, so you can see which section of your site has the problem rather than getting one aggregate number. Image and video sitemaps are handled separately from page sitemaps, because they legitimately repeat page URLs once per asset and would otherwise look like duplicates.
Free, unlimited, measured live from your site — no signup.
Best practices
- Declare the sitemap in robots.txt and submit it in Google Search Console.
- List only canonical, indexable, 200-status URLs.
- Keep lastmod accurate. A file where every date updates nightly is a file Google learns to ignore.
- Split at 50,000 URLs, well before you reach the limit.
- Exclude noindexed pages, redirects, tag archives and paginated series.
- Re-check after any migration or bulk content change.
Common mistakes to avoid
- Listing URLs that redirect, which wastes a crawl request every time.
- Including noindexed pages, which sends contradictory signals.
- Setting every lastmod to today on every regeneration.
- Relying on priority and changefreq — Google ignores both.
- Forgetting to update the sitemap after a domain or protocol migration.
Frequently asked questions
Where should my sitemap live?
Conventionally at /sitemap.xml or /sitemap_index.xml, but the location does not technically matter as long as it is declared in robots.txt and submitted in Search Console. What matters is that crawlers can find it without guessing.
How many URLs can one sitemap hold?
50,000 URLs or 50MB uncompressed, whichever comes first. Beyond that you need multiple sitemaps behind an index file. In practice, splitting well below the limit makes diagnosis much easier because you can see which section has problems.
Do priority and changefreq do anything?
No. Google has stated publicly that it ignores both. lastmod is the only optional field that carries weight, and only when it is accurate — Google learns to distrust files where every date changes on every regeneration.
Should noindexed pages be in the sitemap?
No. A sitemap says “these are my important pages”; a noindex tag says “do not index this page”. Including both on the same URL is a contradiction, and Google resolves it by trusting your sitemap less.
How often is my sitemap recrawled?
It varies by site authority and update frequency — anywhere from hours to weeks. Accurate lastmod dates help significantly, because they let Google prioritize what actually changed instead of recrawling everything.
Do I need an HTML sitemap too?
Rarely, for search purposes. Internal linking does that job better. An HTML sitemap can still be useful for human navigation on large sites, but it is a UX decision rather than an SEO one.
Should my location and service-area pages be in the sitemap?
Yes, and this is worth actually checking rather than assuming. Location pages are frequently the ones a page builder or SEO plugin quietly excludes — they get created outside the normal page flow, or marked noindex during a build and never switched back. A Sarasota business with ten service-area pages and only three of them in the sitemap is a common and entirely invisible problem. Run the check and look for the ones that are missing.
Pick a time. We’ll audit your business. You walk away with a roadmap.
RevSurge Digital is a digital marketing agency in Sarasota, Florida, running technical SEO and answer engine optimization and paid media across Sarasota, Bradenton and the Gulf Coast. Everything this tool flags is work we do every week — grab a free 30-minute audit and we’ll walk you through your report.
Prefer email? Drop us a line at hello@revsurgedigital.com
More tools like this
GEO Audit
Find out whether AI answer engines can actually read, understand and cite your website.
SEO Competitor Analysis
Put your site next to a competitor and see, metric by metric, where you win and where you lose.
Missing Keywords
Find the keywords and topics your site should be targeting and never mentions.
Robots.txt Checker
Fetch and validate any robots.txt in seconds, including your AI crawler policy.
Meta Description Checker
Check any description against Google’s real pixel limits, desktop and mobile.

