Home › Knowledge base › Topics
Block training, allow search
Most providers separate the training crawler from the search crawler. This lets you opt out of model training and still appear as a cited source.
Blocking the training bots
Disallow for: GPTBot (OpenAI), ClaudeBot (Anthropic), CCBot (Common Crawl), plus the tokens Google-Extended and Applebot-Extended. Optionally Meta-ExternalAgent if you also want out of Llama training.
Allowing the search bots
Allow for: OAI-SearchBot and ChatGPT-User (ChatGPT search), Claude-SearchBot and Claude-User (Claude), PerplexityBot and Perplexity-User (Perplexity), Googlebot (Google Search and AI Overviews).
Grey areas
Not every provider separates cleanly: Common Crawl has no search, Bytespider mixes both. And a snippet can end up in an AI Overview even if you have blocked training — that follows search indexing.
Checking the configuration
citeglass shows the applicable robots.txt rule and the real server response for each bot separately — so you can see at a glance whether the "training off, search on" split actually holds.
Check your own site
citeglass shows in about 30 seconds where your website is unreadable for AI systems — free and without login.
Read on
General information, not legal advice. GEO is a young field — recommendations may change.