Robots.txt & AI crawler checker
See which AI crawlers your robots.txt allows or blocks - GPTBot, Google-Extended, PerplexityBot, ClaudeBot, and more - then generate new rules in seconds. Free, no signup.
See which AI crawlers (GPTBot, Google-Extended, PerplexityBot, ClaudeBot, and more) your robots.txt allows or blocks.
Build a robots.txt
Choose which AI crawlers to block, add your sitemap, then copy the result to your site root.
# Allow normal search engines full access User-agent: * Allow: / # Add your sitemap URL, e.g.: # Sitemap: https://example.com/sitemap.xml
What this robots.txt tool does
This tool does two jobs. The checker fetches your site's /robots.txt and tells you, in plain language, which of the major AI crawlers it currently allows or blocks - including OpenAI's GPTBot, Google-Extended, PerplexityBot, ClaudeBot, Common Crawl's CCBot, and more. The generator lets you tick the crawlers you want to block, add your sitemap, and copy a ready-to-deploy robots.txt. It's the fastest way to take control of how AI systems use your content.
How robots.txt works
robots.txt is a plain-text file at your domain root. It follows the Robots Exclusion Protocol. The file is split into groups. Each group starts with one or more User-agent lines (naming a crawler) followed by Allow and Disallow rules. Disallow: / blocks the whole site for that crawler; an empty Disallow: allows everything. A User-agent: * group is the fallback for any crawler you have not named. You can also declare your sitemap with a Sitemap: line. For the full spec, see Google's robots.txt documentation.
One critical caveat: robots.txt is voluntary. Good crawlers honor it. Nothing makes them, and a bad actor can ignore it. To truly block access you need server-side controls. Think of robots.txt as a clear, respected request - not a lock.
The AI crawlers you can control
| Crawler | Vendor | What it's for |
|---|---|---|
| GPTBot | OpenAI | Trains ChatGPT models |
| OAI-SearchBot | OpenAI | Powers ChatGPT search results |
| Google-Extended | Gemini training/grounding (not Search) | |
| PerplexityBot | Perplexity | Indexes pages for Perplexity answers |
| ClaudeBot | Anthropic | Trains and grounds Claude |
| Applebot-Extended | Apple | Apple Intelligence training |
| Bytespider | ByteDance | TikTok/ByteDance AI training |
| CCBot | Common Crawl | Open dataset used to train many LLMs |
Should you block or allow AI crawlers?
This is a strategy decision, not a technical one. There are two distinct kinds of AI crawler, and they deserve different treatment:
- Training crawlers (GPTBot, CCBot, ClaudeBot, Google-Extended, Bytespider) collect content to train models. Block them and your content stays out of training data. That is useful if you sell content, or if you want to guard your IP.
- Search/answer crawlers (OAI-SearchBot, PerplexityBot) fetch pages to cite in live AI answers. Block these and you lose the chance to be cited, plus the referral traffic that comes with it.
Most businesses want visibility. The sane stance is to allow search/answer crawlers (so you get cited and referred). Then decide on training crawlers based on how closely you guard your content. If your growth strategy includes being found in AI answers - as it should in 2026 - blocking the lot is usually a mistake. Pair your decision with our AI search visibility checker and llms.txt generator.
Common robots.txt mistakes
- Accidentally blocking Googlebot. A stray
Disallow: /underUser-agent: *can drop your whole site from the index. Never block a normal search crawler. - Using robots.txt to hide sensitive pages. The file is public, so listing a secret path is a way of advertising it. Put a login in front of it instead.
- Forgetting the sitemap. Add a Sitemap line so crawlers discover all your URLs.
- Wrong location. It must be at the domain root and applies per host/subdomain.
- Confusing crawling with indexing. A Disallow rule stops crawling. It may not stop indexing. Use a noindex meta tag to keep a page out of the results.
Frequently asked questions
- What is robots.txt?
- A plain-text file at your domain root that tells crawlers which parts of your site they may access, using User-agent, Allow, and Disallow rules.
- How do I block GPTBot?
- Add "User-agent: GPTBot" then "Disallow: /". Use the generator above to build rules for any AI crawler and copy them in.
- Should I block AI crawlers?
- It depends. Blocking training crawlers protects content; blocking search crawlers costs you citations and referral traffic. Many allow search crawlers and decide on training crawlers case by case.
- Google-Extended vs Googlebot?
- Googlebot is for normal Search (don't block it). Google-Extended only governs Gemini AI training - blocking it doesn't affect your Search rankings.
- Is this checker free?
- Yes - free, no signup. The checker fetches your public robots.txt; the generator runs in your browser.
- Does robots.txt guarantee a crawler stays out?
- No - it's voluntary. Well-behaved crawlers obey it; to truly restrict access use authentication or firewall rules.
- Where does robots.txt go?
- At the domain root (yourdomain.com/robots.txt). It applies per host and subdomain.
- Should it reference my sitemap?
- Yes - a Sitemap line helps crawlers find all your important URLs. The generator includes it.
Control how AI sees your whole site
- DNS lookup - check the records behind your domain, including the crawler-relevant TXT and NS records.
- llms.txt generator - the emerging companion to robots.txt for guiding AI engines.
- AI search visibility checker - see how ready your site is to be cited by AI.
- On-page SEO checker - confirm pages are indexable and optimized.
- Structured data checker - the schema AI engines rely on to cite you.
For the spec, see RFC 9309 (Robots Exclusion Protocol).
Related free tools
Keep going with these tools
AI Search Visibility Checker
Can ChatGPT and Perplexity cite your site?
AI Content Readiness Scorer
Will AI engines cite your content?
llms.txt Generator
Create a spec-compliant llms.txt file
Schema Markup Generator
Create valid JSON-LD structured data
XML Sitemap Generator
Paste URLs, get a valid sitemap.xml
Hreflang Tag Generator
Correct international SEO tags
Web app development for portals, internal tools, workflows, data products, and SaaS surfaces.
Explore Web app development servicesRelated guides
Go deeper with our guides
Current, practical guides on the strategy, costs, and hands-on work behind this tool.
Technical SEO Checklist for Small Business Websites in 2026
Fix the technical SEO issues holding your site back: crawlability, indexation, site speed, and schema markup.
Small Business SEO Guide 2026: Strategy and Execution
An SEO framework for small businesses in 2026, from technical foundations and content planning to local visibility and tracking.
Content Marketing Strategy for Small Businesses: SEO Guide 2026
Build an SEO-driven content strategy: topic clusters, content calendars, briefs, ROI measurement, and distribution.
Ready to build your next product?
Tell us what you're building. A senior engineer will help you scope it, plan it, and get it built fast, on a foundation that's ready for real users.
