UseToolSuite UseToolSuite

Robots.txt Generator & Validator

Build a robots.txt, then check one you already have: it is parsed the way a crawler parses it, with RFC 9309 group precedence, longest-match Allow/Disallow resolution, and a tester that answers whether a given bot may fetch a given URL.

Configuration

Check a robots.txt you already have

Paste a file to have it parsed the way a crawler parses it — group precedence, longest-match rules, and the directives that quietly do nothing.

Crawl vs index — the distinction that matters

Almost every robots.txt mistake comes from conflating two things:

robots.txtMeta robots / X-Robots-Tag
Controls crawling (fetching)Controls indexing (appearing in results)
Disallow: blocks the bot’s requestnoindex removes the page from results
Page can still be indexed via linksRequires the page to be crawlable to be seen

Internalize this and you’ll use each tool correctly: robots.txt to steer crawlers away from low-value or heavy sections, noindex to actually keep a page out of search.

Anatomy of the file

A robots.txt lives at your domain root (https://example.com/robots.txt — nowhere else) and is grouped by user-agent:

User-agent: *
Disallow: /admin/
Allow: /admin/public/

Sitemap: https://example.com/sitemap.xml

Disallow blocks paths, Allow carves exceptions, and the Sitemap directive (placeable outside any user-agent block) points crawlers to your sitemap so they discover pages that aren’t well-linked.

Don’t block your own CSS and JS

A self-inflicted SEO wound: disallowing your CSS or JavaScript directories. Google renders pages to evaluate them, so if it can’t fetch your stylesheets and scripts, it sees a broken layout and may rank you lower — Search Console will flag this. Always allow crawlers to reach the assets needed to render the page.

It’s advisory, not enforcement

robots.txt is a request, not a lock. Reputable crawlers (Googlebot, Bingbot, GPTBot) obey it; malicious scrapers ignore it entirely. For anything you must truly protect, use authentication or server-level blocking — robots.txt is your first line of guidance, not a security boundary. Pair this with the Privacy Policy Generator and Meta Tag Generator to round out your site’s technical SEO setup.

Robots.txt Generator & Validator Powered by UseToolSuite — free browser tools

Last updated Built and maintained by Necmeddin Cunedioglu How tools are tested

How helpful was this tool?

Click to rate

Key Concepts

Robots Exclusion Protocol

A standard (originally proposed in 1994) that defines how web crawlers should interact with websites. The robots.txt file, the noindex meta tag, and the X-Robots-Tag HTTP header are all part of this protocol. Compliance is voluntary — well-behaved crawlers follow the rules, but the protocol provides no enforcement mechanism.

User-agent Directive

A robots.txt directive that specifies which crawler the following rules apply to. User-agent: * matches all crawlers, while User-agent: Googlebot targets only Google's crawler. Multiple User-agent blocks can be defined in one robots.txt file to apply different rules to different bots.

Crawl-delay Directive

A robots.txt directive that requests crawlers to wait a specified number of seconds between requests. For example, Crawl-delay: 10 asks bots to wait 10 seconds between each page fetch. Bing respects this directive; Google does not — use Google Search Console's crawl rate settings instead. Use crawl-delay when your server has limited resources and heavy bot traffic is causing performance issues.

Sitemap Directive

A robots.txt directive that declares the location of your XML sitemap: Sitemap: https://example.com/sitemap.xml. This helps search engines discover all pages on your site, including those not reachable through internal links. Multiple Sitemap directives can be specified. The Sitemap directive is not tied to any User-agent block and applies to all crawlers.

Frequently Asked Questions

Where should I place the robots.txt file?

The robots.txt file must be placed at the root of your domain: https://example.com/robots.txt. It must be accessible at that exact URL. Placing it in a subdirectory will not work — crawlers only look for it at the domain root. The file must be served with a text/plain Content-Type header.

Does robots.txt prevent pages from appearing in search results?

No. Robots.txt blocks crawling, not indexing. If other pages link to a URL that is disallowed in robots.txt, search engines may still index the URL based on anchor text and link context — they just won't crawl the page content. To prevent indexing, use the noindex meta tag or X-Robots-Tag HTTP header instead.

How do I block AI crawlers like GPTBot and Google-Extended?

Add separate User-agent blocks for each AI crawler you want to block: User-agent: GPTBot followed by Disallow: /. This tool includes presets for major AI crawlers including GPTBot (OpenAI), ChatGPT-User, Google-Extended (Gemini), CCBot (Common Crawl), anthropic-ai, and Bytespider (TikTok). Note that compliance is voluntary — well-behaved bots respect robots.txt, but malicious scrapers may ignore it.

Can I use wildcards in robots.txt paths?

Yes. Google and Bing support the * wildcard and the $ end-of-URL anchor. For example, Disallow: /*.pdf$ blocks all URLs ending in .pdf, and Disallow: /*/private/ blocks any URL containing /private/ in the path. However, not all crawlers support these extensions — the original robots.txt specification only defines exact prefix matching.

What is the Sitemap directive and why is it important?

The Sitemap directive tells crawlers where to find your XML sitemap: Sitemap: https://example.com/sitemap.xml. This is important because it helps search engines discover pages that might not be reachable through internal links alone. The Sitemap directive can be placed outside any User-agent block and applies globally.

Why do my Disallow rules not apply to Googlebot?

Because Googlebot has its own group. Under RFC 9309 a crawler obeys only the single most specific group that names it, and ignores every other group completely — including "User-agent: *". So a file that says "User-agent: Googlebot / Allow: /" to welcome Google, and then puts the real restrictions under "User-agent: *", has exempted Googlebot from all of them. Either repeat the rules inside the named group, or list the named agents alongside * in one group. The validator on this page flags exactly this pattern.

If both an Allow and a Disallow match a URL, which wins?

The more specific rule, measured by the number of characters in the path pattern. "Disallow: /folder/" and "Allow: /folder/public/" both match /folder/public/page, but the Allow pattern is longer, so the URL is crawlable. When two patterns are exactly the same length, Allow wins. The tester on this page shows which rule won and why.

Does Crawl-delay do anything?

Not for Google — it has never supported the directive and ignores it silently. Bing and Yandex do honour it. If you need Googlebot to slow down, use the crawl rate setting in Search Console, or return 503/429 responses. The validator warns when it sees Crawl-delay so you do not assume it is working.

Will robots.txt remove my page from Google search?

No — this is the single most common robots.txt misconception. robots.txt controls CRAWLING (whether bots fetch the page), not INDEXING (whether it appears in search results). If another site links to a URL you've disallowed, Google can still INDEX that URL and show it in results — just with no description, since it couldn't crawl the content. So disallowing a page you want hidden can backfire: it stays in the index AND Google can't see the noindex tag that would actually remove it. To truly keep a page out of search, ALLOW crawling and add <meta name='robots' content='noindex'> (or an X-Robots-Tag: noindex header) so Google can read the directive. Use robots.txt to manage crawl budget and block bots from heavy/irrelevant paths — not as a privacy or de-indexing tool.

How do I block AI training crawlers like GPTBot?

Add a separate User-agent block for each AI crawler with Disallow: /. The major ones to name: GPTBot (OpenAI training), ChatGPT-User (ChatGPT browsing), Google-Extended (Gemini/Bard training, does NOT affect Google Search ranking), CCBot (Common Crawl, feeds many AI datasets), anthropic-ai / ClaudeBot (Anthropic), PerplexityBot, and Bytespider (TikTok). For example: 'User-agent: GPTBot' then 'Disallow: /'. This generator includes presets for these. Two caveats: compliance is VOLUNTARY — well-behaved bots honor it, but you can't force a scraper that ignores robots.txt, so for hard enforcement also block their user-agent strings at your server/CDN; and blocking Google-Extended opts you out of AI training without hurting your normal Google Search visibility, which many sites want.

Troubleshooting & Technical Tips

Pages still appearing in Google despite Disallow rule

Robots.txt blocks crawling, not indexing. If external sites link to your disallowed pages, Google can still index the URL (showing a title and snippet derived from links, not page content). To fully prevent indexing, add a <meta name="robots" content="noindex"> tag to the page HTML, or send an X-Robots-Tag: noindex HTTP header. Note: Google must be able to crawl a page to see a noindex tag, so do not both disallow and noindex the same URL.

AI bots ignoring robots.txt: GPTBot still scraping content

Robots.txt compliance is voluntary. While reputable bots (Googlebot, Bingbot, GPTBot) honor robots.txt, some AI scrapers do not. For additional protection: use your web server or CDN (Cloudflare, Fastly) to block known AI bot user-agent strings at the server level, implement rate limiting, or require JavaScript rendering that simple scrapers cannot execute. Robots.txt should be your first line of defense but not your only one.

Googlebot cannot access CSS and JS files: Page renders incorrectly

If you disallow CSS or JavaScript directories in robots.txt, Googlebot cannot render your pages properly — leading to poor indexing and potential ranking drops. Google's Search Console will flag this as a crawling issue. Always allow access to CSS, JS, and image files that are needed for page rendering. Use Allow: /wp-content/uploads/ and Allow: /wp-includes/css/ in WordPress robots.txt configurations.

Related Guides

Related Tools

Embed this tool on your site

Paste this snippet into any HTML page or blog post to embed a live, fully working copy of Robots.txt Generator & Validator. Free for any use.