English
citeglass · AI visibility check

HomeKnowledge base › Topics

Block training, allow search

Most providers separate the training crawler from the search crawler. This lets you opt out of model training and still appear as a cited source.

Blocking the training bots

Disallow for: GPTBot (OpenAI), ClaudeBot (Anthropic), CCBot (Common Crawl), plus the tokens Google-Extended and Applebot-Extended. Optionally Meta-ExternalAgent if you also want out of Llama training.

Allowing the search bots

Allow for: OAI-SearchBot and ChatGPT-User (ChatGPT search), Claude-SearchBot and Claude-User (Claude), PerplexityBot and Perplexity-User (Perplexity), Googlebot (Google Search and AI Overviews).

Grey areas

Not every provider separates cleanly: Common Crawl has no search, Bytespider mixes both. And a snippet can end up in an AI Overview even if you have blocked training — that follows search indexing.

Checking the configuration

citeglass shows the applicable robots.txt rule and the real server response for each bot separately — so you can see at a glance whether the "training off, search on" split actually holds.

Check your own site

citeglass shows in about 30 seconds where your website is unreadable for AI systems — free and without login.

Read on

General information, not legal advice. GEO is a young field — recommendations may change.