AI Crawlers Directory

Every AI crawler we check — who runs it, what it does, and whether you should allow it. 40 entries across three jobs. Each profile links to the operator's own documentation where one exists, with the date we last verified it.

Retrieval / search crawlers 12

Build the index AI answers cite from. Blocking these removes you from AI answers — the expensive mistake.

OAI-SearchBot

OpenAI's search crawler. Sites opted out of it are not shown in ChatGPT search answers — allow it if you want citations.

OAI-AdsBot

Visits landing pages submitted as ChatGPT ads for policy review. Only those pages, and the data is not used to train foundation models.

Claude-SearchBot

Anthropic's search crawler; disabling it may reduce your visibility and accuracy in Claude search results.

PerplexityBot

Perplexity's search crawler. Perplexity recommends allowing it plus its published IP ranges; it is not used to train foundation models.

Googlebot

Google's main search crawler; controls Google Search, including AI features built on the Search index. Almost never block.

Bingbot

Bing index and Microsoft Copilot retrieval; almost never block.

Applebot

Powers Spotlight, Siri and Safari search. Training use is governed by the separate Applebot-Extended token.

DuckAssistBot

DuckDuckGo AI summaries.

meta-webindexer

Meta AI search index; blocking removes you from Meta AI answers.

PetalBot

Petal Search / Celia assistant index.

ImagesiftBot

Hive's AI image search index.

iaskspider

iAsk.ai AI search engine index.

User-triggered fetchers 6

Fetch a page in real time when a user opens your link in a chat. Safe to allow — note that some operators say robots.txt may not apply to these.

ChatGPT-User

Fetches a page when someone opens a link inside ChatGPT. OpenAI states robots.txt rules may not apply because the request is user-triggered.

Claude-User

User-initiated fetch for Claude; disabling it may reduce visibility in user-directed search.

Perplexity-User

User-initiated fetch. Perplexity states this fetcher generally ignores robots.txt rules; it is not a crawler and is not used for training.

MistralAI-User

User-initiated fetch when Le Chat users open links.

Meta-ExternalFetcher

User-initiated fetch when Meta AI users open links.

DeepSeek-User

User-initiated fetch when DeepSeek users open links.

Training / data crawlers 22

Harvest content to train models or build datasets. Blocking opts you out of training — a legitimate choice. Some entries here are robots.txt control tokens, not crawlers.

GPTBot

Trains OpenAI models; blocking opts your content out of training. It does not by itself control ChatGPT citations — that is OAI-SearchBot.

ClaudeBot

Trains Claude models; restricting it signals your content should be excluded from Anthropic training datasets.

anthropic-ai

Legacy Anthropic training token, superseded by ClaudeBot. Still found in older robots.txt blocks.

Claude-Web

Deprecated Anthropic crawler token. Harmless to leave blocked, but it is no longer the token Anthropic documents.

Google-Extended

Not a crawler — a robots.txt control token with no separate HTTP user-agent. Controls whether crawled content may be used for Gemini training and Gemini grounding. Google states it does not affect Search inclusion or ranking.

GoogleOther

Google's generic crawler used for one-off research and internal R&D fetches. Google states it does not affect any specific product.

Applebot-Extended

Not a crawler — a control token that decides how Apple may use content already crawled by Applebot. Opts you out of Apple foundation-model training.

CCBot

Feeds the Common Crawl open dataset, which many model trainers build on.

Bytespider

ByteDance data collection. Frequently reported to ignore robots.txt — community-reported; ByteDance publishes limited crawler documentation.

Meta-ExternalAgent

Meta AI training and data collection.

FacebookBot

Meta's older crawler token, largely superseded by Meta-ExternalAgent.

Amazonbot

Amazon describes it as powering Alexa answers; grouped here as data collection, though Amazon's own wording is closer to retrieval.

cohere-ai

Cohere model training crawler.

cohere-training-data-crawler

Cohere legacy training-data crawler token.

Diffbot

Commercial knowledge-graph and data-broker crawler; blocking prevents content resale.

MistralAI-Training

Mistral model training crawler.

YouBot

You.com index and training crawler.

Webzio-Extended

Commercial data broker licensing content for AI training.

AI2Bot

Academic research crawler (Allen AI); builds open datasets such as Dolma.

PanguBot

Huawei Pangu model training.

Timpibot

Timpi independent web index (search and AI data).

DeepSeek-Bot

DeepSeek model training and data collection.

Check your site against all 40 →

checkaibots.com — AI visibility checker · Free · No signup