AI Crawlers and robots.txt: GPTBot, OAI-SearchBot, ClaudeBot and More

Which AI crawlers visit your site, what each one is for and how to allow AI search while blocking AI training in robots.txt. With copy-paste examples.

AI crawlers and robots.txt: AI search crawlers allowed, AI training crawlers blocked

AI companies send different crawlers for different jobs, and robots.txt can treat each one differently. The difference that matters most is between training crawlers, which collect pages to train models, and search crawlers, which fetch pages so an assistant can cite them in an answer. You can block training and still appear in ChatGPT, Claude and Perplexity answers, provided the search crawlers are allowed.

The three kinds of AI crawlers

Three kinds of AI crawlers: training crawlers, AI search crawlers and user-triggered fetchers
Training, search and user fetches are separate bots

Training crawlers download pages to build datasets for future models. Blocking them doesn't remove you from AI answers today.

Search crawlers build the index an assistant searches when it answers with sources. If you block them, you disappear from those answers.

User-triggered fetchers visit a page because a person asked the assistant to open it. Some vendors say robots.txt may not apply to these, since a human started the visit.

The user agents that matter

CompanyUser agentWhat it doesBlocking it means
OpenAIGPTBotCollects pages for model trainingNot used to train future OpenAI models
OpenAIOAI-SearchBotIndexes pages for ChatGPT searchNot shown or cited in ChatGPT search answers
OpenAIChatGPT-UserOpens pages when a user asks ChatGPT toUser-initiated; OpenAI notes robots.txt may not apply
AnthropicClaudeBotCollects pages for model trainingNot used to train Claude
AnthropicClaude-SearchBotIndexes pages for Claude's search resultsLess visible in Claude answers with web search
AnthropicClaude-UserOpens pages when a user asks Claude toClaude can't fetch your page for that user
PerplexityPerplexityBotIndexes pages for Perplexity answers (not training)Not cited in Perplexity answers
PerplexityPerplexity-UserOpens pages when a user asksUser-initiated; Perplexity says it generally ignores robots.txt
GoogleGooglebotGoogle Search, including AI Overviews and AI ModeOut of Google Search entirely
GoogleGoogle-ExtendedA robots.txt token, not a crawler: controls use of your content for Gemini training and groundingDoesn't affect Search or AI Overviews
AppleApplebot-ExtendedA token that controls use for Apple's model trainingApplebot still crawls for Siri and Spotlight
MetaMeta-ExternalAgentCollects pages for Meta's AI trainingNot used to train Meta's models
Common CrawlCCBotOpen web archive widely used to train modelsOut of future Common Crawl snapshots
ByteDanceBytespiderCollects pages for ByteDance's products, reportedly including AI trainingOut of its crawl if it obeys robots.txt. Reports say it doesn't always, so many sites also block it at the CDN

The table follows each vendor's own documentation: OpenAI, Anthropic, Perplexity, Google and Apple. Vendors add and rename bots from time to time, so check those pages before relying on any list, this one included.

Google works differently. There is no separate AI crawler for AI Overviews or AI Mode, which use the normal Googlebot index. Google-Extended only controls Gemini apps and model training. To limit what Google can quote, use nosnippet, max-snippet or data-nosnippet on the page.

Example robots.txt files

Allow AI search, block AI training

Most businesses that want to be cited, but don't want their pages used as training data, end up with something like this:

# AI search and user fetches: allowed
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
Allow: /

# AI training: blocked
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: Meta-ExternalAgent
User-agent: CCBot
User-agent: Bytespider
Disallow: /

# Everyone else, including Googlebot and Bingbot
User-agent: *
Allow: /

Sitemap: https://www.example.com/sitemap.xml

Allow everything

If you want the widest possible reach in AI answers and don't mind training use, you don't need any AI-specific rules:

User-agent: *
Allow: /

Sitemap: https://www.example.com/sitemap.xml

Keep a private section out of every bot

User-agent: *
Disallow: /account/
Disallow: /checkout/

robots.txt only asks bots to stay away and doesn't stop anyone from opening a page. Anything private needs a login.

Common mistakes

  • Allowing Googlebot and Bingbot and putting a blanket Disallow: / on everything else. That quietly blocks every AI search crawler too.
  • Bot protection at the CDN or firewall. It can block AI crawlers even when robots.txt allows them, so check your CDN's AI-bot settings as well.
  • Grouping rules wrongly. A crawler follows only the most specific User-agent group that matches it (Google's robots.txt rules explain the matching), so once you name GPTBot in a group, the rules under User-agent: * stop applying to it.
  • Content that only appears after JavaScript runs. Most AI crawlers don't run scripts, so even when they're allowed in, they see an empty page.
  • Blocking search crawlers to stop training. Blocking OAI-SearchBot or PerplexityBot does nothing about training and only removes you from answers.

How to check your site

  1. Open yourdomain.com/robots.txt and look for the user agents above.
  2. Run the free robots.txt checker. It tests Googlebot and the main AI crawlers against your rules for any page and shows which line decides.
  3. Check your server or CDN logs for the user agents to confirm they actually reach you.
  4. Make sure your main pages work without JavaScript. The website audit flags pages whose content only appears after scripts run.

FAQ

Should I block GPTBot?

It depends on whether you mind your content being used to train OpenAI's models. Blocking GPTBot doesn't remove you from ChatGPT search answers, which depend on OAI-SearchBot. Many publishers block GPTBot and allow OAI-SearchBot.

Does blocking Google-Extended remove me from AI Overviews?

No. AI Overviews and AI Mode are part of Google Search and use Googlebot. Google-Extended only controls whether content is used for Gemini apps and model training. Use nosnippet or max-snippet to limit what Google quotes.

Do AI crawlers respect robots.txt?

OpenAI, Anthropic, Google, Apple and Perplexity (for PerplexityBot) say their crawlers follow robots.txt. User-triggered fetchers are a gray area, and some smaller crawlers ignore it. For hard blocking, use your firewall or CDN.

What is llms.txt and does it replace robots.txt?

No. llms.txt is a proposed Markdown file that gives AI tools a short map of your site. It doesn't grant or deny access; robots.txt still controls crawling. What is llms.txt explains the format, and the llms.txt validator checks yours.