Every time a search engine visits your site, it first looks for a small text file at your domain root called robots.txt. This file tells crawlers which parts of your site they may visit and which to leave alone — admin panels, cart pages, staging areas and internal search results. One wrong line can accidentally block your entire site from Google, which is why correct syntax matters.
The format itself is simple: blocks of User-agent lines followed by Allow and Disallow rules, with an optional Sitemap line pointing to your XML sitemap. Simple does not mean foolproof, though — a stray slash or a misplaced wildcard changes the meaning of a rule completely.
This free robots.txt generator builds the file for you with a visual rule editor. Add allow and disallow rules per user-agent, attach your sitemap URL, set a crawl delay if you need one, then download a ready-to-upload robots.txt.
How to use the robots.txt generator
- Review the starter rules. The tool begins with a sensible default: everything allowed except /admin. Edit or remove these to match your site.
- Add your rules. Click “Add rule” for each path you want to control. Set the user-agent (
*means all crawlers, or target specific ones like Googlebot), choose Allow or Disallow, and enter the path starting with a slash. - Add your sitemap URL. Paste the full URL of your XML sitemap so crawlers discover it immediately — this one line meaningfully speeds up indexing of new sites.
- Set a crawl delay if needed. Most sites should leave this blank. Only add one if a specific crawler is hammering a fragile server.
- Copy or download the file. Use the Copy button for quick pastes or Download to get a robots.txt file ready for upload.
- Upload to your domain root. Place the file at https://yourdomain.com/robots.txt — it only works at the root, not in a subdirectory. Then verify it loads in your browser.
Key features and benefits
- Visual rule builder. No memorizing syntax — pick the user-agent, the rule type and the path from guided fields, and the file assembles itself correctly.
- Multiple user-agent blocks. Create different rules for all crawlers (
*) versus specific bots, grouped properly in the output so each block is valid. - Sitemap line included. One field adds the
Sitemap:directive pointing crawlers at your sitemap — the fastest way to get new pages discovered. - Optional crawl delay. Add a polite pause between crawler requests for bots that respect it, without cluttering the file when you do not need it.
- Path validation. The download step checks that your paths start with a forward slash and warns you before you ship a broken file.
- One-click download. Get a properly named robots.txt file instantly — no renaming, no encoding issues, ready to upload via FTP or your hosting panel.
- Free and private. Your site structure never leaves your browser. Plan rules for unreleased projects without exposing anything.
Robots.txt best practices that prevent disasters
Most robots.txt horror stories share one trait: someone blocked more than they intended. A few habits keep you safe.
Never disallow what you want indexed
The classic catastrophe is Disallow: / left over from a staging site, which tells every crawler to ignore your entire domain. Before uploading, read the generated file line by line and ask of each rule: “do I really want this hidden from Google?” When in doubt, leave it allowed — you can always restrict later.
Block utilities, not content
Good candidates for disallowing: admin areas, login pages, cart and checkout flows, internal site-search result pages, and API endpoints. Bad candidates: anything with content you want found — blog posts, product pages, category archives. Remember that robots.txt is a suggestion, not security: it does not hide pages from determined visitors, so never rely on it to protect sensitive data.
Keep it minimal and test it
Shorter files are easier to audit. After uploading, check https://yourdomain.com/robots.txt in your browser to confirm it served correctly, then use Google Search Console’s robots.txt tester to verify Google interprets your rules the way you intended. Revisit the file whenever you launch new site sections.
Path Syntax: Slashes and Special Characters
Robots.txt looks simple until a rule silently fails. Every path must start with a forward slash, wildcards (* and $) have specific meanings, and special characters in URLs need proper encoding — a literal space or unencoded character can make a rule match nothing. The visual builder here removes the memorization: pick the user-agent, choose Allow or Disallow, type the path, and the file assembles with valid syntax. Paths with query strings or odd characters? Clean them first with the URL encoder-decoder so the rule matches what you intend.
The Sitemap Line Crawlers Actually Read
The single highest-value line in a robots.txt is often the Sitemap: directive. It points crawlers straight at your XML sitemap, which meaningfully speeds up discovery and indexing of new pages — especially on new sites with few backlinks. Add your full sitemap URL in the builder’s dedicated field and it lands in the file automatically. Keep the sitemap itself clean and valid; if the XML needs tidying, run it through the XML formatter first. Then upload the finished file to your domain root and verify it loads in a browser.
Frequently asked questions
Where do I upload my robots.txt file?
It must live at the root of your domain: https://yourdomain.com/robots.txt. Crawlers will not look for it in subdirectories. Upload it via FTP, your hosting file manager, or your CMS — WordPress users can also generate one through SEO plugins, but a hand-built file gives you full control.
What does “User-agent: *” mean?
The asterisk is a wildcard meaning “all crawlers”. Rules under User-agent: * apply to Googlebot, Bingbot and every other bot unless a more specific block overrides them. Use specific user-agent names like Googlebot only when you want different rules for different crawlers.
Can robots.txt keep a page completely out of Google?
Not reliably. Disallow stops crawlers from visiting, but Google can still index a URL it discovers through links — sometimes showing it without a description. If a page must stay out of search results, use a noindex robots meta tag (which you can add with the meta tag generator) instead of, or in addition to, a robots.txt rule.
Should I add a crawl delay?
Almost certainly not. Googlebot ignores crawl-delay entirely, and most healthy servers handle normal crawling fine. Only consider it if a specific bot is causing measurable server load — server-level rate limiting is usually the better fix.
How does the sitemap line help?
The Sitemap: directive tells every crawler exactly where your XML sitemap lives, without you having to submit it to each search engine manually. It is especially valuable for new sites with few inbound links, where crawlers might otherwise take weeks to discover all your pages.
Is robots.txt a security tool?
No — and this is critical to understand. Robots.txt is publicly readable by anyone, so listing /secret-admin-panel in it actually advertises that path to attackers. It only asks polite crawlers to stay out. Real protection for sensitive areas means authentication, not robots.txt.
Where do I upload the robots.txt file?
Your domain root — https://yourdomain.com/robots.txt. It only works at the root, not in a subdirectory.
Does robots.txt improve my rankings?
No. It guides crawlers; it does not boost rankings. Blocking the wrong paths, though, can hurt — so validate every path.
What is crawl delay?
A polite pause between a crawler’s requests. Most sites should leave it blank unless a bot is hammering a fragile server.
How It Works: Under the Hood
A robots.txt file is a plain-text contract between your site and crawlers, built from four directives. User-agent names the bot the rules apply to — * for all bots, or specific names like Googlebot. Disallow lists paths the bot must not crawl, Allow carves exceptions back out, and Sitemap points to your XML sitemap.
Path matching uses simple substring rules, not regex. * matches any characters, so Disallow: /*? blocks every URL with a query string; $ anchors a match to the URL’s end, so Disallow: /*.php$ blocks .php URLs but allows /file.php?v=2. When Allow and Disallow both match, the longer path wins. A single Disallow: / blocks everything — the root prefixes every URL — and is the most common way sites accidentally deindex themselves.
Crucially, robots.txt controls crawling, not indexing — a disallowed page can still appear in results if other sites link to it. That differs from the noindex meta tag, which the crawler must fetch the page to read, meaning you must allow crawling for noindex to work. Doing both simultaneously is a contradiction the crawler never resolves.
Real-World Use Cases
- A store owner blocks /cart/, /checkout/, and /wp-admin/ so Google spends crawl budget on product pages instead of thin utility pages.
- An SEO specialist auditing a site that vanished from Google finds a leftover staging Disallow: / and generates a corrected file, restoring crawling.
- A WordPress blogger blocks /?s= and /tag/ archives that created thousands of near-duplicate indexed pages.
- A developer opens /docs/v2/ to crawlers with an Allow exception while keeping /docs/ blocked, so only the new version gets indexed.
- A news publisher adds a sitemap reference so fresh articles are discovered within minutes of publishing.
Advanced Tips
- Stick to one User-agent: * group unless a specific bot needs different treatment — extra groups multiply maintenance and are easy to misread.
- Test before deploying. Google Search Console’s robots.txt tester shows exactly which URLs your rules match — run your most important pages through it before uploading.
- Keep the file under 500 KB — Google stops reading larger files. If your rule list is enormous, use server-side blocking instead.
- Carve exceptions with Allow. Disallow: /private/ plus Allow: /private/public-report.pdf works because the longer match wins — the cleanest way to open one file inside a blocked directory.
Common Mistakes to Avoid
- Pushing Disallow: / live. The number one robots.txt disaster — usually a staging config copied to production. Verify the root path is not disallowed before uploading.
- Expecting Disallow to remove pages from Google. It does not — disallowed pages can still be indexed via links. To remove one, noindex it and allow the crawl so Google sees the tag.
- Disallowing CSS and JavaScript. Google needs to render your pages — blocking asset folders makes them render as broken text to Googlebot and hurts rankings.
- Treating robots.txt as security. Well-behaved crawlers obey it; aggressive scrapers do not — sensitive content needs real authentication, not a Disallow line.