The most damaging robots.txt mistake is not blocking too much. It is blocking a page you also want deindexed, which achieves the opposite.
Crawling is not indexing
These are separate stages and conflating them causes the classic failure.
Crawling is fetching the page. robots.txt controls this.
Indexing is deciding to store it and show it in results. Meta robots tags and HTTP headers control this.
If you block a page in robots.txt and it has inbound links, Google may index it anyway from those links alone — producing a result with the URL and no description, sometimes with "No information is available for this page". That is worse than either leaving it alone or deindexing it properly.
The mistake, stated plainly
To remove a page from search, Google must be able to crawl it and find the noindex instruction. Blocking it in robots.txt prevents exactly that.
- Allow crawling of the page in robots.txt.
- Add
<meta name="robots" content="noindex,follow">to the page, or the equivalent X-Robots-Tag header. - Wait for it to be recrawled and dropped.
- Only then, if you want to save crawl budget, add the robots.txt block.
Doing steps 1 and 4 in the wrong order is the single most common technical SEO error, and it is self-perpetuating — the page stays indexed, so someone adds a stronger block, which keeps it indexed.
What robots.txt is genuinely for
- Saving crawl budget on large sites — faceted search URLs, infinite calendars, parameter permutations that generate millions of near-identical pages.
- Keeping crawlers out of expensive endpoints — internal search, export scripts, anything that costs CPU per request.
- Pointing at your sitemap with a
Sitemap:line, which is the conventional place for it. - Controlling which bots may crawl at all, including AI crawlers such as GPTBot and ClaudeBot.
That last one is worth a deliberate decision rather than a default. Blocking AI crawlers while hoping to be cited by AI assistants is a contradiction, and it is one people arrive at by copying a robots.txt from somewhere else.
Rules that catch people out
| Rule | Effect |
|---|---|
Disallow: (empty) | Allows everything |
Disallow: / | Blocks the entire site |
Disallow: /admin | Blocks /admin and /administrator too |
Disallow: /admin/ | Blocks only the directory |
| Case | Paths are case-sensitive |
| Order | Most specific match wins, not first match |
The prefix behaviour in row three is the one that surprises. Disallow: /admin matches any path starting with those characters, so it silently blocks /administration-guide as well. Add the trailing slash unless you mean the prefix.
What a sitemap does and does not do
A sitemap lists URLs you would like discovered. It is a hint, not an instruction — inclusion never forces indexing, and Google routinely ignores lastmod values it finds unreliable.
Its real value is on large sites, new sites with few inbound links, and pages that are poorly linked internally. On a small well-linked site it adds very little, because the crawler would find everything anyway.
- Only include canonical, indexable URLs. Listing a page you have noindexed sends contradictory signals.
- 50,000 URLs or 50 MB per file, then split and use a sitemap index. This site emits several by section.
- Accurate lastmod, or omit it. Setting every page to today on every build teaches Google to ignore the field.
- Skip priority and changefreq. Google has said publicly it ignores both.
- Reference it from robots.txt and submit it in Search Console.
Generating the sitemap from the same data that builds the pages is what stops it listing URLs that no longer exist — the XML Sitemap Generator handles the format, and the discipline is to generate rather than maintain.
Checking it
Three checks worth running after any change, because a robots.txt mistake is silent and expensive.
- Fetch it. Visit
yoursite.com/robots.txtand read what is actually served, which is not always what you deployed. - Test specific URLs in Search Console's robots.txt tester against the pages you care about most.
- Check Coverage in Search Console for "Indexed, though blocked by robots.txt" — that report is exactly the failure described above, and it names the affected pages.
A wrong robots.txt can take a site out of search entirely and produce no error anywhere. It is worth checking after every deployment that touches it, which is a habit rather than a task.
Frequently asked questions
Does robots.txt remove a page from Google?
No. It stops crawling, not indexing. A blocked page with inbound links can still appear in results — as a bare URL with no description, since Google cannot see what is on it. To remove a page, allow crawling and use a noindex tag.
How do I properly remove a page from search?
Allow crawling, add a noindex meta tag or X-Robots-Tag header, and wait for a recrawl. Only add a robots.txt block afterwards, if at all — blocking first prevents Google from ever seeing the noindex instruction.
Why is Disallow: /admin blocking other pages?
Because it matches any path starting with those characters, so /administration-guide is blocked too. Add a trailing slash — Disallow: /admin/ — unless you genuinely mean the prefix.
Does a sitemap improve rankings?
No. It aids discovery, which matters on large sites, new sites with few links, and poorly linked pages. On a small well-linked site the crawler would find everything anyway. Inclusion never forces indexing.
Should I set priority and changefreq in my sitemap?
No — Google has said publicly it ignores both. Accurate lastmod is worth including; setting every page to today on every build teaches Google to disregard that field too, so omit it if you cannot make it truthful.
Should I block AI crawlers in robots.txt?
It is a deliberate decision, not a default. Blocking GPTBot or ClaudeBot while hoping to be cited by AI assistants is a contradiction — and it is one people reach by copying someone else's robots.txt without reading it.