Every AI crawler we check — who runs it, what it does, and whether you should allow it. 40 entries across three jobs. Each profile links to the operator's own documentation where one exists, with the date we last verified it.
Build the index AI answers cite from. Blocking these removes you from AI answers — the expensive mistake.
OpenAI's search crawler. Sites opted out of it are not shown in ChatGPT search answers — allow it if you want citations.
Visits landing pages submitted as ChatGPT ads for policy review. Only those pages, and the data is not used to train foundation models.
Anthropic's search crawler; disabling it may reduce your visibility and accuracy in Claude search results.
Perplexity's search crawler. Perplexity recommends allowing it plus its published IP ranges; it is not used to train foundation models.
Google's main search crawler; controls Google Search, including AI features built on the Search index. Almost never block.
Bing index and Microsoft Copilot retrieval; almost never block.
Powers Spotlight, Siri and Safari search. Training use is governed by the separate Applebot-Extended token.
DuckDuckGo AI summaries.
Meta AI search index; blocking removes you from Meta AI answers.
Petal Search / Celia assistant index.
Hive's AI image search index.
iAsk.ai AI search engine index.
Fetch a page in real time when a user opens your link in a chat. Safe to allow — note that some operators say robots.txt may not apply to these.
Fetches a page when someone opens a link inside ChatGPT. OpenAI states robots.txt rules may not apply because the request is user-triggered.
User-initiated fetch for Claude; disabling it may reduce visibility in user-directed search.
User-initiated fetch. Perplexity states this fetcher generally ignores robots.txt rules; it is not a crawler and is not used for training.
User-initiated fetch when Le Chat users open links.
User-initiated fetch when Meta AI users open links.
User-initiated fetch when DeepSeek users open links.
Harvest content to train models or build datasets. Blocking opts you out of training — a legitimate choice. Some entries here are robots.txt control tokens, not crawlers.
Trains OpenAI models; blocking opts your content out of training. It does not by itself control ChatGPT citations — that is OAI-SearchBot.
Trains Claude models; restricting it signals your content should be excluded from Anthropic training datasets.
Legacy Anthropic training token, superseded by ClaudeBot. Still found in older robots.txt blocks.
Deprecated Anthropic crawler token. Harmless to leave blocked, but it is no longer the token Anthropic documents.
Not a crawler — a robots.txt control token with no separate HTTP user-agent. Controls whether crawled content may be used for Gemini training and Gemini grounding. Google states it does not affect Search inclusion or ranking.
Google's generic crawler used for one-off research and internal R&D fetches. Google states it does not affect any specific product.
Not a crawler — a control token that decides how Apple may use content already crawled by Applebot. Opts you out of Apple foundation-model training.
Feeds the Common Crawl open dataset, which many model trainers build on.
ByteDance data collection. Frequently reported to ignore robots.txt — community-reported; ByteDance publishes limited crawler documentation.
Meta AI training and data collection.
Meta's older crawler token, largely superseded by Meta-ExternalAgent.
Amazon describes it as powering Alexa answers; grouped here as data collection, though Amazon's own wording is closer to retrieval.
Cohere model training crawler.
Cohere legacy training-data crawler token.
Commercial knowledge-graph and data-broker crawler; blocking prevents content resale.
Mistral model training crawler.
You.com index and training crawler.
Commercial data broker licensing content for AI training.
Academic research crawler (Allen AI); builds open datasets such as Dolma.
Huawei Pangu model training.
Timpi independent web index (search and AI data).
DeepSeek model training and data collection.
Check your site against all 40 →