What is a robots.txt file?
A robots.txt file is a plain text file at the root of your site, for example https://example.com/robots.txt, that tells crawlers which paths they may fetch. It is made of groups. Each group starts with one or more User-agent lines that name a crawler, followed by Allow and Disallow rules. Sitemap lines point crawlers to your XML sitemaps and can sit anywhere in the file. The format is standardised as RFC 9309, the Robots Exclusion Protocol, and Google, Bing and most large crawlers follow it.
This robots.txt generator writes the file for you from presets and a simple group editor. The robots.txt validator flags lines that crawlers will ignore or misread, with line numbers. The robots.txt tester tells you whether a URL is allowed or blocked for a given bot and shows the exact line that decided it.
What robots.txt can and cannot do
-
It controls crawling, not indexing. A blocked URL can still appear in Google results, without a description, if other pages link to it. To keep a page out of search, let crawlers fetch it and add a
noindexrobots meta tag or anX-Robots-Tagheader. - It is voluntary. Well-behaved crawlers follow it. Scrapers that ignore it are not stopped. It is not access control: protect private pages with a login.
- It is public. Anyone can open your robots.txt, so do not list secret paths in it.
-
It works per host.
shop.example.comandexample.comeach need their own file at their own root.
How Google reads the rules
Google documents its matching logic, and the Test URL tab follows the same steps:
-
Pick the group. A crawler uses the group with the most specific user-agent that matches its token, case-insensitively. Googlebot-News uses a
googlebot-newsgroup if there is one, otherwisegooglebot, otherwise*. Several groups for the same crawler are merged into one. -
Match the path. Paths are case-sensitive and match from the start, including the query string.
*matches any run of characters and$marks the end of the URL./fish*means the same as/fish, and/*.php$matches/index.phpbut not/index.php?id=1. - The longest rule wins. Among matching rules, the one with the longest path decides. If an Allow and a Disallow rule are equally long, Google uses the least restrictive one, which is the Allow.
-
An empty Disallow blocks nothing.
Disallow:with no path means the group may crawl everything.
User-agent: * Disallow: /account/ Allow: /account/login
Here /account/orders is blocked, while /account/login is allowed because the Allow rule is longer and more specific.
Blocking AI crawlers with robots.txt
Many AI companies publish user-agent tokens you can block. They are not all the same kind of bot, and the difference matters:
- Training crawlers. GPTBot (OpenAI), ClaudeBot (Anthropic), CCBot (Common Crawl), Meta-ExternalAgent and Bytespider (ByteDance) collect pages that can be used to train models. Blocking them asks these companies not to use future crawls.
- Control tokens. Google-Extended and Applebot-Extended are not separate crawlers. Google-Extended tells Google not to use your content for Gemini training and grounding. It does not change how Googlebot crawls, indexes or ranks your pages in Google Search. Applebot-Extended opts you out of Apple's model training while Applebot keeps crawling for Siri and Spotlight.
- AI search and user fetches. OAI-SearchBot, Claude-SearchBot and PerplexityBot fetch pages so they can show and cite them in AI answers. ChatGPT-User and Claude-User fetch a page when a person asks the assistant to read it. Blocking these can remove your pages from AI answers and citations, which is often traffic you want.
The AI preset blocks training crawlers and leaves search and user agents unchecked so you can decide. robots.txt only applies to future crawls. It does not remove content that was already collected.
robots.txt on Shopify and WordPress
Shopify generates robots.txt for you and you cannot upload a file. To change it, add a robots.txt.liquid template to your theme (Online Store -> Themes -> Edit code -> Add a new template -> robots) and add your rules there. The Shopify-style preset in this tool is similar to Shopify's default, so you can review and test rules before you edit the template. WordPress serves a virtual robots.txt that disallows /wp-admin/ but allows /wp-admin/admin-ajax.php. A real robots.txt file in the site root replaces it, and many SEO plugins let you edit it from the dashboard.
Common mistakes the validator catches
-
Noindexlines. Google stopped supporting them in robots.txt in 2019 and ignores them. - Relative sitemap URLs such as
Sitemap: /sitemap.xml. The URL must be absolute. - Rules placed before any
User-agentline, which crawlers ignore. - Paths without a leading slash, full URLs in rules, and non-ASCII paths that should be percent-encoded.
-
Crawl-delay, which Google ignores. Bing and Yandex read it. - Files larger than 500 KiB. Google ignores everything after that limit.
- Typos such as
Dissallowand unknown directives.
Crawl rules are only one part of technical SEO. Check your structured data with the schema markup validator, write search snippets with the meta description generator, run a page through the AI SEO analyzer, or browse all free AI generators and tools.
FAQ
Is this robots.txt generator free?
Yes. There is no sign-up and no limit. Everything runs in your browser, so the robots.txt you build, paste or test is not sent to any server.
Where do I put the robots.txt file?
In the root of each host, so it loads at https://yourdomain.com/robots.txt. A robots.txt in a subfolder is ignored. On Shopify you edit the robots.txt.liquid theme template instead of uploading a file.
Does robots.txt stop a page from being indexed?
No. It controls crawling, not indexing. A blocked page can still be indexed from links, just without its content. To keep a page out of Google, allow crawling and add a noindex robots meta tag or an X-Robots-Tag header.
Will blocking GPTBot or Google-Extended hurt my Google rankings?
No. GPTBot is OpenAI's training crawler and has nothing to do with Google Search. Google-Extended only controls whether Google uses your content for Gemini training and grounding, not crawling or ranking in Search. Blocking AI search bots such as OAI-SearchBot or PerplexityBot can remove your pages from those AI answers, though.
Do all bots obey robots.txt?
No. robots.txt is a voluntary standard. Major search engines and most AI companies say they follow it, but scrapers can ignore it. Use logins, firewall rules or rate limits for anything that must stay private.