All tools run in your browser — your files never leave your device.
All tools154

Design & SEO

Robots.txt & llms.txt Generator

Build robots.txt and llms.txt with presets for every major crawler.

What it does. robots.txt tells crawlers which parts of your site they may request. It lives at your domain root and is the first file most crawlers fetch. This generator builds valid rules with presets for search and AI crawlers, and also produces llms.txt — the emerging convention for describing a site to language models.
Runs in your browserNothing uploadsNo signupWorks offline

How to use Robots.txt & llms.txt Generator

  1. Choose whether to allow or block each crawler group, and add any paths to disallow.
  2. Add your sitemap URL and, optionally, a crawl delay.
  3. Copy the robots.txt output to your site root, and the llms.txt output alongside it.

Which crawlers should I allow?

The default should be allowing everything. Blocking crawlers is a decision with real costs and it is made too casually.

Which crawlers should I allow?
User-agentOperatorRecommendation
GooglebotGoogle SearchAlways allow
BingbotBing, and the index behind CopilotAlways allow
Google-ExtendedGemini training and groundingAllow if you want AI citations
GPTBotOpenAI trainingAllow for visibility in ChatGPT
OAI-SearchBotChatGPT search resultsAllow — this is search traffic
ClaudeBotAnthropicAllow for citations in Claude
PerplexityBotPerplexityAllow — Perplexity links out prominently
CCBotCommon CrawlAllow unless you object to bulk archiving
AhrefsBot / SemrushBotSEO toolingBlock if crawl budget is a concern

The AI crawler question is a genuine strategic decision rather than a technical one. Blocking them protects content from training use; it also removes you from the answers those systems give, which is an increasingly significant traffic source. For a site whose business is being found, blocking is usually the wrong trade.

What robots.txt does not do

It does not remove pages from search results. This is the most consequential misunderstanding about the file.

Disallow prevents crawling, not indexing. If a disallowed URL is linked from elsewhere, Google can and does index it using only the anchor text — which produces a result with a title, no description, and the note "no information is available for this page". The page is in the index, and blocking it in robots.txt is what prevented Google from seeing the noindex tag that would have removed it.

To remove a page from results, allow crawling and serve <meta name="robots" content="noindex"> or an X-Robots-Tag header. To keep a page private, require authentication. robots.txt is a publicly readable request, not an access control — listing /admin/ in it tells everyone where your admin panel is.

What is llms.txt?

A proposed convention, introduced in 2024, for a Markdown file at your site root that describes your site to language models in a form they can consume directly.

The reasoning is that an AI crawler reading a modern web page has to strip navigation, ads, cookie banners and scripts to find the content. llms.txt skips that: a short description of what the site is, followed by a curated list of your most important URLs with one-line summaries.

It is not yet a standard and no major AI provider has committed to reading it. It is cheap to publish — this generator produces one from your page list — and the downside is a file nobody fetches. Given how quickly AI-mediated discovery is growing, that is a reasonable bet.

Common robots.txt mistakes

  • Blocking CSS and JavaScript. Google renders pages to evaluate them. Blocking assets means it sees a broken layout and may judge the page unusable on mobile.
  • Using robots.txt to hide sensitive paths. The file is public. Anyone can read it, and listing private directories advertises them.
  • Leaving a staging Disallow: / in production. A single line shipped from a staging environment can deindex an entire site. It happens regularly.
  • Expecting Disallow to remove indexed pages. It prevents crawling, which prevents the noindex tag from being seen.
  • Wrong location. It must be at the domain root. A file at /blog/robots.txt is ignored entirely, and subdomains need their own.
  • Assuming Crawl-delay works for Google. Googlebot ignores it. Bing and Yandex honor it. Use Search Console to adjust Google’s crawl rate.

Frequently asked questions

Where does robots.txt go?

The root of the domain — example.com/robots.txt. It is ignored anywhere else, and each subdomain needs its own file.

Does robots.txt remove a page from Google?

No. It prevents crawling. A disallowed URL linked from elsewhere can still appear in results without a description. Use a noindex meta tag instead, and allow crawling so it can be seen.

Should I block AI crawlers?

It is a trade. Blocking protects content from training use and removes you from AI-generated answers, which is a growing traffic source. Most sites that want to be found should allow them.

Is robots.txt legally binding?

No. It is a voluntary convention. Well-behaved crawlers honor it; scrapers ignore it. Use authentication for anything that must not be accessed.

Do I need a robots.txt at all?

Not strictly. A missing file returns 404 and crawlers proceed as if everything is allowed. Having one lets you point to your sitemap and set explicit rules.

What is the difference between Disallow and noindex?

Disallow stops the crawler fetching the page. noindex lets it fetch and tells it not to index. They are frequently confused, and using Disallow when you meant noindex leaves the page indexed.

Can I test my robots.txt?

Yes, with the robots.txt report in Google Search Console, which shows exactly how Googlebot parses your rules.

Guides for Robots.txt & llms.txt Generator