Skip to content

Sitemap URL extractor

Paste an XML sitemap, get one URL per line — deduplicated, counted, and ready to paste into whatever needs a list.

Runs entirely in your browser — nothing is uploaded, and it is free with no account.

A sitemap is a claim, not a result

It is worth being clear about what the list you just extracted represents. A sitemap is the set of URLs you say are worth crawling. It is an invitation, and Google treats it as one: it reads the file, decides which URLs to fetch and when, and then decides separately whether each fetched page earns a place in the index. Both decisions can go against you, and neither is reported in the sitemap itself.

That gap is what the Search Console statuses describe. Discovered – currently not indexed means Google read the URL here and never came to fetch it. Crawled – currently not indexed means it came, read the page, and declined. A sitemap listing a thousand URLs tells you nothing about which of the two happened, or to how many.

The one thing a sitemap does uniquely well is lastmod. It is the only signal you control that tells Google a page changed, and it is only believed if it is accurate — a file that stamps today's date on every URL at each deploy gets the field ignored entirely.

One thing worth doing with the list once you have it: check that every URL returns a 200 rather than a redirect. A sitemap should name destinations, and a redirecting entry in one is the single case where "Page with redirect" in Search Console is worth acting on rather than ignoring.

Questions

What is a sitemap index, and why did I get no URLs?
A sitemap index is a sitemap of sitemaps: instead of <url> entries it holds <sitemap> entries pointing at other files. Large sites use one because a single sitemap is capped at 50,000 URLs or 50MB. If you pasted an index, this tool lists the child sitemaps for you — open each one and paste it in turn.
Can it fetch the sitemap from a URL?
No, and nothing that runs purely in your browser can. The same-origin policy stops a page reading a file from another domain, so any tool offering that is running a server which fetches on your behalf. Open the sitemap URL in a tab and copy the XML — your browser is already allowed to read it.
Does being in a sitemap mean a page is indexed?
No. A sitemap is a request, not a guarantee. Google treats it as a hint about which URLs exist and when they changed, then decides independently whether each one is worth indexing — which is why "Discovered – currently not indexed" exists as a status.
What does it do with duplicates?
Removes them, and tells you how many it removed. Duplicates are common when a sitemap is generated by more than one plugin, and they cost you real money in any tool that charges per URL.

Checking whether those URLs are actually in Google?

That is the part a browser cannot do for you. indexaction runs up to 1,000 URLs in one check and returns a verdict for each — no site: searches, no CAPTCHA. New accounts get 50 free credits.

Check 50 URLs free

Other free tools

indexaction on Facebook Email [email protected]