AI Crawlers and robots.txt: GPTBot, OAI-SearchBot, ClaudeBot and More
Which AI crawlers visit your site, what each one is for and how to allow AI search while blocking AI training in robots.txt. With copy-paste examples.
AI companies send different crawlers for different jobs, and robots.txt can treat each one differently. The difference that matters most is between training crawlers, which collect pages to train models, and search crawlers, which fetch pages so an assistant can cite them in an answer. You can block training and still appear in ChatGPT, Claude and Perplexity answers, provided the search crawlers are allowed.
The three kinds of AI crawlers
Training crawlers download pages to build datasets for future models. Blocking them doesn't remove you from AI answers today.
Search crawlers build the index an assistant searches when it answers with sources. If you block them, you disappear from those answers.
User-triggered fetchers visit a page because a person asked the assistant to open it. Some vendors say robots.txt may not apply to these, since a human started the visit.
The user agents that matter
| Company | User agent | What it does | Blocking it means |
|---|---|---|---|
| OpenAI | GPTBot | Collects pages for model training | Not used to train future OpenAI models |
| OpenAI | OAI-SearchBot | Indexes pages for ChatGPT search | Not shown or cited in ChatGPT search answers |
| OpenAI | ChatGPT-User | Opens pages when a user asks ChatGPT to | User-initiated; OpenAI notes robots.txt may not apply |
| Anthropic | ClaudeBot | Collects pages for model training | Not used to train Claude |
| Anthropic | Claude-SearchBot | Indexes pages for Claude's search results | Less visible in Claude answers with web search |
| Anthropic | Claude-User | Opens pages when a user asks Claude to | Claude can't fetch your page for that user |
| Perplexity | PerplexityBot | Indexes pages for Perplexity answers (not training) | Not cited in Perplexity answers |
| Perplexity | Perplexity-User | Opens pages when a user asks | User-initiated; Perplexity says it generally ignores robots.txt |
Googlebot | Google Search, including AI Overviews and AI Mode | Out of Google Search entirely | |
Google-Extended | A robots.txt token, not a crawler: controls use of your content for Gemini training and grounding | Doesn't affect Search or AI Overviews | |
| Apple | Applebot-Extended | A token that controls use for Apple's model training | Applebot still crawls for Siri and Spotlight |
| Meta | Meta-ExternalAgent | Collects pages for Meta's AI training | Not used to train Meta's models |
| Common Crawl | CCBot | Open web archive widely used to train models | Out of future Common Crawl snapshots |
| ByteDance | Bytespider | Collects pages for ByteDance's products, reportedly including AI training | Out of its crawl if it obeys robots.txt. Reports say it doesn't always, so many sites also block it at the CDN |
The table follows each vendor's own documentation: OpenAI, Anthropic, Perplexity, Google and Apple. Vendors add and rename bots from time to time, so check those pages before relying on any list, this one included.
Google works differently. There is no separate AI crawler for AI Overviews or AI Mode, which use the normal Googlebot index.
Google-Extendedonly controls Gemini apps and model training. To limit what Google can quote, usenosnippet,max-snippetordata-nosnippeton the page.
Example robots.txt files
Allow AI search, block AI training
Most businesses that want to be cited, but don't want their pages used as training data, end up with something like this:
# AI search and user fetches: allowed
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
Allow: /
# AI training: blocked
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: Meta-ExternalAgent
User-agent: CCBot
User-agent: Bytespider
Disallow: /
# Everyone else, including Googlebot and Bingbot
User-agent: *
Allow: /
Sitemap: https://www.example.com/sitemap.xml
Allow everything
If you want the widest possible reach in AI answers and don't mind training use, you don't need any AI-specific rules:
User-agent: *
Allow: /
Sitemap: https://www.example.com/sitemap.xml
Keep a private section out of every bot
User-agent: *
Disallow: /account/
Disallow: /checkout/
robots.txt only asks bots to stay away and doesn't stop anyone from opening a page. Anything private needs a login.
Common mistakes
- Allowing Googlebot and Bingbot and putting a blanket
Disallow: /on everything else. That quietly blocks every AI search crawler too. - Bot protection at the CDN or firewall. It can block AI crawlers even when robots.txt allows them, so check your CDN's AI-bot settings as well.
- Grouping rules wrongly. A crawler follows only the most specific
User-agentgroup that matches it (Google's robots.txt rules explain the matching), so once you nameGPTBotin a group, the rules underUser-agent: *stop applying to it. - Content that only appears after JavaScript runs. Most AI crawlers don't run scripts, so even when they're allowed in, they see an empty page.
- Blocking search crawlers to stop training. Blocking
OAI-SearchBotorPerplexityBotdoes nothing about training and only removes you from answers.
How to check your site
- Open
yourdomain.com/robots.txtand look for the user agents above. - Run the free robots.txt checker. It tests Googlebot and the main AI crawlers against your rules for any page and shows which line decides.
- Check your server or CDN logs for the user agents to confirm they actually reach you.
- Make sure your main pages work without JavaScript. The website audit flags pages whose content only appears after scripts run.
FAQ
Should I block GPTBot?
It depends on whether you mind your content being used to train OpenAI's models. Blocking GPTBot doesn't remove you from ChatGPT search answers, which depend on OAI-SearchBot. Many publishers block GPTBot and allow OAI-SearchBot.
Does blocking Google-Extended remove me from AI Overviews?
No. AI Overviews and AI Mode are part of Google Search and use Googlebot. Google-Extended only controls whether content is used for Gemini apps and model training. Use nosnippet or max-snippet to limit what Google quotes.
Do AI crawlers respect robots.txt?
OpenAI, Anthropic, Google, Apple and Perplexity (for PerplexityBot) say their crawlers follow robots.txt. User-triggered fetchers are a gray area, and some smaller crawlers ignore it. For hard blocking, use your firewall or CDN.
What is llms.txt and does it replace robots.txt?
No. llms.txt is a proposed Markdown file that gives AI tools a short map of your site. It doesn't grant or deny access; robots.txt still controls crawling. What is llms.txt explains the format, and the llms.txt validator checks yours.