Rankite
ServicesResultsToolsTeamAboutBlogCareersContactFree SEO Audit
Free tool

AI Crawler robots.txt Generator

Pick which AI bots to block, add the generated block to your existing robots.txt, and control exactly who trains on or reads your site.

Home / Tools / AI Crawler robots.txt Generator

Check any bot below to add a block instructing it not to crawl your site. Leave a bot unchecked to allow it. Paste the result near the top of your existing robots.txt file.

Add this to your robots.txt

Check a bot above to build its Disallow block.

This only affects the bots you check. Regular search crawlers like Googlebot and Bingbot are never touched by this tool.

Built by Rankite, the SEO team behind Swordfish AI's +400% revenue and Zluri's +45% organic growth. See the case studies

Search engines are not the only thing crawling your site anymore. A growing set of AI-specific bots visit to collect training data, power live citations in tools like ChatGPT and Perplexity, or both, and each one identifies itself with its own user agent string in robots.txt. This tool lets you decide, bot by bot, which ones get in, without touching Googlebot, Bingbot or anything else already governing your normal search visibility.

Training crawlers vs on-demand fetchers

These bots fall into two different categories with very different implications. Training crawlers, GPTBot, ClaudeBot, CCBot, Bytespider and similar, visit on their own schedule to collect content that may end up shaping a future model's training data. On-demand fetchers, OAI-SearchBot, PerplexityBot, ChatGPT-User and similar, instead retrieve a specific page live, usually because a real user's question needs an answer from it right now, functioning much closer to how a search engine crawler works today. Blocking the first group protects your content from training use, blocking the second can reduce your odds of being cited as a source in an AI-generated answer.

Google-Extended and Applebot-Extended are different

Two of the entries on this list are not crawlers at all. Google-Extended and Applebot-Extended make no HTTP requests of their own, they exist purely as robots.txt tokens you can add to opt specific AI training uses out, for Gemini and Apple Intelligence respectively, without affecting Googlebot's or Applebot's normal search crawling in any way. Checking one of these only changes an AI training permission, your regular search presence stays exactly as it was.

Where to put this block

Add the generated lines near the top of your existing robots.txt, before your general User-agent: * rules, then verify the file is still reachable at yourdomain.com/robots.txt. This tool generates a focused, AI-specific addition, if you also need a full, general-purpose robots.txt from scratch, use the robots.txt generator linked below instead.

Related articles

FAQ

AI Crawler robots.txt Generator: questions, answered

Will blocking these crawlers remove my site from Google Search or ChatGPT answers?
No, as long as you leave Googlebot and the standard search crawlers alone. The bots listed here are separate, AI-specific user agents. Blocking GPTBot, for example, stops OpenAI from using your pages to train future models, it does not touch Google Search indexing or Googlebot at all.
What is the difference between Google-Extended and Googlebot?
Googlebot crawls and indexes your site for regular Google Search and is not affected by anything in this tool. Google-Extended is a separate, opt-out-only token: it makes no crawl requests of its own, it is simply a signal you can add to your robots.txt to tell Google not to use your content to train Gemini and other AI models. The same relationship holds between Applebot, which powers Siri and Spotlight, and Applebot-Extended, its AI-training opt-out.
Do AI companies actually respect robots.txt?
OpenAI, Anthropic, Google, Apple and Perplexity all publicly state that their listed crawlers honor robots.txt Disallow rules. Bytespider and some smaller or less transparent crawlers have a mixed compliance record, so treat robots.txt as a clear, industry-standard signal of your preference rather than an unbreakable technical lock.
Should I block AI crawlers or allow them?
There is no single right answer. Blocking training crawlers like GPTBot and ClaudeBot keeps your content out of future model training data, which some publishers want for control or licensing reasons. Allowing on-demand fetchers like OAI-SearchBot or PerplexityBot can help your pages get cited as sources in AI answers, which is closer to how organic search visibility works today. Many sites block training bots while leaving the answer-engine fetchers alone.
Does blocking GPTBot stop my content from ever appearing in ChatGPT?
Not necessarily. GPTBot is OpenAI's training crawler. ChatGPT's browsing and search features use a separate user agent, OAI-SearchBot for indexing and ChatGPT-User for on-demand fetches when a user asks it to check a page live, and those are controlled independently in this tool.
How do I know if these bots are already crawling my site?
Check your server access logs or a tool like Cloudflare's bot analytics for the exact user agent strings listed on this page, GPTBot, ClaudeBot, PerplexityBot and so on. Seeing hits from a bot confirms it is currently allowed, either because your robots.txt has no rule for it or because it is choosing to ignore one.

More free tools

Let's grow

Ready to own page one?

Get a free, no-obligation SEO audit and a 30-minute strategy session. We'll show you exactly where the growth is hiding.

Book your free audit Explore services
Get in touch

Tell us about your project

Fill out the form and we'll get back to you within one business day. Prefer email? Write to us directly at contact@rankite.com.

Or copy our email and write to us directly: contact@rankite.com