// free tool - robots.txt generator

Robots.txt Generator Validate and test it too

Build a robots.txt file from presets for Shopify, WordPress and AI crawlers, check it line by line for mistakes, and test whether a URL is allowed or blocked for Googlebot, GPTBot or any other bot, using Google's matching rules.

  • Free - no sign-up
  • Runs in your browser
  • AI crawler presets
  • Google matching rules

Your robots.txt

Nothing is sent anywhere. The output updates as you type.

Start from a preset

User-agent groups
AI crawlers to block (0 selected)

Checked bots get their own group with Disallow: /. robots.txt is voluntary: well-behaved bots follow it, scrapers that ignore it are not stopped.

    robots.txt Ready
    Checks

      Same URL, other bots

      User-agent Result Decided by
      Go past the file

      A clean robots.txt is a start. Technical SEO is the rest.

      Crawl rules are one part of how search engines and AI assistants see your site. These AI specialists plug into Claude, ChatGPT or any AI chat and help with the rest: audits, indexing problems, site structure and content.

      What is a robots.txt file?

      A robots.txt file is a plain text file at the root of your site, for example https://example.com/robots.txt, that tells crawlers which paths they may fetch. It is made of groups. Each group starts with one or more User-agent lines that name a crawler, followed by Allow and Disallow rules. Sitemap lines point crawlers to your XML sitemaps and can sit anywhere in the file. The format is standardised as RFC 9309, the Robots Exclusion Protocol, and Google, Bing and most large crawlers follow it.

      This robots.txt generator writes the file for you from presets and a simple group editor. The robots.txt validator flags lines that crawlers will ignore or misread, with line numbers. The robots.txt tester tells you whether a URL is allowed or blocked for a given bot and shows the exact line that decided it.

      What robots.txt can and cannot do

      • It controls crawling, not indexing. A blocked URL can still appear in Google results, without a description, if other pages link to it. To keep a page out of search, let crawlers fetch it and add a noindex robots meta tag or an X-Robots-Tag header.
      • It is voluntary. Well-behaved crawlers follow it. Scrapers that ignore it are not stopped. It is not access control: protect private pages with a login.
      • It is public. Anyone can open your robots.txt, so do not list secret paths in it.
      • It works per host. shop.example.com and example.com each need their own file at their own root.

      How Google reads the rules

      Google documents its matching logic, and the Test URL tab follows the same steps:

      1. Pick the group. A crawler uses the group with the most specific user-agent that matches its token, case-insensitively. Googlebot-News uses a googlebot-news group if there is one, otherwise googlebot, otherwise *. Several groups for the same crawler are merged into one.
      2. Match the path. Paths are case-sensitive and match from the start, including the query string. * matches any run of characters and $ marks the end of the URL. /fish* means the same as /fish, and /*.php$ matches /index.php but not /index.php?id=1.
      3. The longest rule wins. Among matching rules, the one with the longest path decides. If an Allow and a Disallow rule are equally long, Google uses the least restrictive one, which is the Allow.
      4. An empty Disallow blocks nothing. Disallow: with no path means the group may crawl everything.
      User-agent: *
      Disallow: /account/
      Allow: /account/login

      Here /account/orders is blocked, while /account/login is allowed because the Allow rule is longer and more specific.

      Blocking AI crawlers with robots.txt

      Many AI companies publish user-agent tokens you can block. They are not all the same kind of bot, and the difference matters:

      • Training crawlers. GPTBot (OpenAI), ClaudeBot (Anthropic), CCBot (Common Crawl), Meta-ExternalAgent and Bytespider (ByteDance) collect pages that can be used to train models. Blocking them asks these companies not to use future crawls.
      • Control tokens. Google-Extended and Applebot-Extended are not separate crawlers. Google-Extended tells Google not to use your content for Gemini training and grounding. It does not change how Googlebot crawls, indexes or ranks your pages in Google Search. Applebot-Extended opts you out of Apple's model training while Applebot keeps crawling for Siri and Spotlight.
      • AI search and user fetches. OAI-SearchBot, Claude-SearchBot and PerplexityBot fetch pages so they can show and cite them in AI answers. ChatGPT-User and Claude-User fetch a page when a person asks the assistant to read it. Blocking these can remove your pages from AI answers and citations, which is often traffic you want.

      The AI preset blocks training crawlers and leaves search and user agents unchecked so you can decide. robots.txt only applies to future crawls. It does not remove content that was already collected.

      robots.txt on Shopify and WordPress

      Shopify generates robots.txt for you and you cannot upload a file. To change it, add a robots.txt.liquid template to your theme (Online Store -> Themes -> Edit code -> Add a new template -> robots) and add your rules there. The Shopify-style preset in this tool is similar to Shopify's default, so you can review and test rules before you edit the template. WordPress serves a virtual robots.txt that disallows /wp-admin/ but allows /wp-admin/admin-ajax.php. A real robots.txt file in the site root replaces it, and many SEO plugins let you edit it from the dashboard.

      Common mistakes the validator catches

      • Noindex lines. Google stopped supporting them in robots.txt in 2019 and ignores them.
      • Relative sitemap URLs such as Sitemap: /sitemap.xml. The URL must be absolute.
      • Rules placed before any User-agent line, which crawlers ignore.
      • Paths without a leading slash, full URLs in rules, and non-ASCII paths that should be percent-encoded.
      • Crawl-delay, which Google ignores. Bing and Yandex read it.
      • Files larger than 500 KiB. Google ignores everything after that limit.
      • Typos such as Dissallow and unknown directives.

      Crawl rules are only one part of technical SEO. Check your structured data with the schema markup validator, write search snippets with the meta description generator, run a page through the AI SEO analyzer, or browse all free AI generators and tools.

      FAQ

      Is this robots.txt generator free?

      Yes. There is no sign-up and no limit. Everything runs in your browser, so the robots.txt you build, paste or test is not sent to any server.

      Where do I put the robots.txt file?

      In the root of each host, so it loads at https://yourdomain.com/robots.txt. A robots.txt in a subfolder is ignored. On Shopify you edit the robots.txt.liquid theme template instead of uploading a file.

      Does robots.txt stop a page from being indexed?

      No. It controls crawling, not indexing. A blocked page can still be indexed from links, just without its content. To keep a page out of Google, allow crawling and add a noindex robots meta tag or an X-Robots-Tag header.

      Will blocking GPTBot or Google-Extended hurt my Google rankings?

      No. GPTBot is OpenAI's training crawler and has nothing to do with Google Search. Google-Extended only controls whether Google uses your content for Gemini training and grounding, not crawling or ranking in Search. Blocking AI search bots such as OAI-SearchBot or PerplexityBot can remove your pages from those AI answers, though.

      Do all bots obey robots.txt?

      No. robots.txt is a voluntary standard. Major search engines and most AI companies say they follow it, but scrapers can ignore it. Use logins, firewall rules or rate limits for anything that must stay private.