AI crawler directory
The AI crawlers your robots.txt decides for
These are the 14 AI crawler tokens CrawlRadar reads your robots.txt for, one group at a time, for the exact path scanned — and only 3 of them (OAI-SearchBot, Claude-SearchBot, PerplexityBot) build the search indexes answer engines cite from. The rest train models or fetch pages on request.
Check my robots.txt- 14 tokens checked
- 3 search crawlers
By Iliass El Barhoumi · Published · Last updated
AI crawlers fall into four groups, and only one of them decides citations. Search crawlers build the indexes answer engines quote from; live fetchers open a page because someone asked; training crawlers and policy tokens decide whether your pages feed future models.
Search crawlers: they decide citations
Blocking one removes you from that engine’s answers. CrawlRadar flags a block as critical.
Live fetchers: a person asked for your page
Each request is a conversation that needed your page. Blocking them costs the highest-intent visits.
Training crawlers: an opt-out, not a blocker
Blocking them keeps your pages out of model training without removing you from any answer engine.
- CrawlRadar checks GPTBot. OpenAI: OpenAI model training
GPTBot
- CrawlRadar checks ClaudeBot. Anthropic: Anthropic model training
ClaudeBot
- CrawlRadar checks anthropic-ai. Anthropic: Anthropic (legacy token)
anthropic-ai
- CCBotCrawlRadar checks CCBot. Common Crawl: Common Crawl (feeds many models)
- CrawlRadar checks Bytespider. ByteDance: ByteDance model training
Bytespider
- CrawlRadar checks Amazonbot. Amazon: Amazon (Alexa answers, may train models)
Amazonbot
Policy tokens: read from robots.txt, never sent
No request carries these names. They control how the vendor’s regular crawl may be used.
Frequently asked questions
Which AI crawlers should I allow in robots.txt?
Allow the search crawlers, OAI-SearchBot, Claude-SearchBot and PerplexityBot, if you want ChatGPT, Claude and Perplexity to cite you. Allowing or blocking the training crawlers, such as GPTBot, ClaudeBot and CCBot, is a separate choice about model training that does not affect citations.
Is it safe to block GPTBot?
Yes, if you do not want your pages used to train OpenAI’s models. OpenAI documents GPTBot as a training crawler and OAI-SearchBot as the one behind ChatGPT Search, and states the two settings are independent.
Does Google-Extended control Google AI Overviews?
No. Google-Extended governs Gemini training and grounding, and Google states it does not affect inclusion in Google Search. AI Overviews are a Search feature built from Googlebot’s crawl.