How to block AI training but allow AI search in robots.txt
Many sites want to be found and quoted in AI search, but not used to train AI models. AI companies use separate crawlers for those jobs, so robots.txt can allow one and block the other.
Which crawler does what
- Training: GPTBot (OpenAI), ClaudeBot (Anthropic), Google-Extended (Gemini), Applebot-Extended (Apple Intelligence), CCBot (Common Crawl, used by many models), Meta-ExternalAgent (Meta).
- AI search index: OAI-SearchBot (ChatGPT search), Claude-SearchBot (Claude), PerplexityBot (Perplexity).
- Fetches when a user asks: ChatGPT-User, Claude-User, Perplexity-User.
Example robots.txt
This keeps regular search and AI search open and blocks training. Keep your existing rules for other crawlers and your Sitemap line above it.
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: Applebot-Extended
Disallow: /
User-agent: CCBot
Disallow: /
User-agent: Meta-ExternalAgent
Disallow: /
User-agent: OAI-SearchBot
Allow: /
User-agent: Claude-SearchBot
Allow: /
User-agent: PerplexityBot
Allow: /Google-Extended does not affect Google Search: Googlebot keeps crawling and ranking your pages. robots.txt is a request that well-known crawlers say they follow, not a technical lock.
Check it before and after
Paste your new file into the robots.txt tester to confirm each crawler gets the answer you expect, then test the live file again after you upload it.
Questions
Will this hurt my Google rankings?
No. Googlebot is not affected by these rules. Only Google-Extended is blocked, which controls use of your content for Gemini.
Is blocking training retroactive?
No. It stops future crawling. Content collected before the change may already be in existing training data.