robots.txt tester
Paste a robots.txt and a URL. You get a verdict and, more usefully, the exact line that produced it.
Runs entirely in your browser — nothing is uploaded, and it is free with no account.
How Google actually reads this file
Three rules explain almost every surprise. The first is that order does not matter. Google picks the matching rule with the longest path pattern and applies it, wherever it sits in the file, and only falls back to "allow wins" when two matching rules are exactly the same length. A file that reads top-to-bottom like a sequence of instructions is not being read that way.
The second is that a crawler obeys exactly one group. If there is a
User-agent: Googlebot block, Googlebot reads it and ignores the
User-agent: * block entirely — including any Disallow you assumed
applied to everyone. Rules are not inherited from the wildcard.
One group, though, can be assembled from several blocks naming the same
crawler. Two User-agent: GPTBot blocks are read as one set of rules, so
a Disallow: / in the first and an Allow: / in the second is a
same-length tie, and allow wins. That is not a hypothetical: a CDN that manages robots.txt
for you prepends its own blocks above yours, and the file ends up saying both things. This
tester merges them, and tells you how many blocks it merged.
The third is the one that costs real traffic: Disallow is not noindex.
Blocking a URL stops Google fetching it, which also stops Google seeing any
noindex on the page. If other sites link to it, the URL can sit in the index
indefinitely with no description — the "Indexed, though blocked by robots.txt" status. To
remove a page you have to let Google crawl it and read the tag telling it to go away.
If a page you expected to see in Google is missing, the causes are worth walking in order — why is my page not indexed by Google covers all six, robots.txt being only the first.
Questions
- Which rule wins when Allow and Disallow both match?
- The more specific one, measured by the length of the path pattern — not the order they appear in the file. Disallow: / and Allow: /blog/ means /blog/post is allowed, because /blog/ is six characters and / is one. When two matching rules are exactly the same length, Allow wins.
- Does blocking a URL in robots.txt remove it from Google?
- No, and this is the single most expensive misunderstanding about robots.txt. Disallow stops Google crawling the page, which also stops it seeing a noindex tag on that page. A blocked URL with links pointing at it can stay in the index indefinitely, listed without a description. To remove a page, let Google crawl it and serve noindex.
- Why does this ask me to paste the file instead of fetching it?
- A browser cannot read another site’s robots.txt — the same-origin policy blocks it, and any tool that does fetch it is running a server that does so on your behalf. Pasting keeps this page free, instant, and free of anything to abuse. Open https://example.com/robots.txt and copy what you see.
- What happens when the same crawler appears in two blocks?
- They are merged into one group, and then the ordinary precedence rules run over the combined rules. So Disallow: / in the first GPTBot block and Allow: / in the second is a tie on path length, and allow wins — the crawler is allowed. This comes up whenever a CDN manages robots.txt for you: Cloudflare, for one, prepends its own Disallow: / for GPTBot, ClaudeBot, CCBot, Google-Extended and Applebot-Extended above whatever your site serves. Two blocks for one agent looks like a mistake in the file and usually is not.
- Does it handle wildcards?
- Yes. * matches any sequence of characters and $ anchors the match to the end of the URL path, both as Google implements them. So Disallow: /*.pdf$ blocks /report.pdf but not /report.pdf.html.
Checking whether those URLs are actually in Google?
That is the part a browser cannot do for you. indexaction runs up to 1,000 URLs in one
check and returns a verdict for each — no site: searches, no CAPTCHA. New
accounts get 50 free credits.
Other free tools
- Sitemap URL extractor — Turn a sitemap into a plain list of URLs you can paste anywhere.
- SERP snippet preview — See where Google cuts your title and description, measured in pixels.
- Bulk URL cleaner — Deduplicate and normalise a messy list of links before you check it.