# checkaibots.com

> Free AI crawler visibility checker. Enter a URL to see which of 40 AI crawlers (GPTBot, ClaudeBot, PerplexityBot, Google-Extended and more) can access a site based on its robots.txt — including the exact blocking rule, how to fix it, a UA-level HTTP fetch check for GPTBot / OAI-SearchBot / PerplexityBot, and free daily monitoring with change alerts. Every crawler profile links to the operator's own documentation where one exists, with the date we last verified it.

## Tools

- [AI Crawler Checker](https://checkaibots.com/): instant free check of 40 AI crawlers against any site, with the exact robots.txt rule behind each result.
- [AI Crawlers Directory](https://checkaibots.com/crawlers): reference profiles for every AI crawler we check — operator, type, and whether you should allow it.

## Crawler guides

- [GPTBot](https://checkaibots.com/crawler/gptbot): OpenAI's training/data crawler. Trains OpenAI models; blocking opts your content out of training. It does not by itself control ChatGPT citations — that is OAI-SearchBot.
- [ClaudeBot](https://checkaibots.com/crawler/claudebot): Anthropic's training/data crawler. Trains Claude models; restricting it signals your content should be excluded from Anthropic training datasets.
- [anthropic-ai](https://checkaibots.com/crawler/anthropic-ai): Anthropic's training/data crawler. Legacy Anthropic training token, superseded by ClaudeBot. Still found in older robots.txt blocks.
- [Claude-Web](https://checkaibots.com/crawler/claude-web): Anthropic's training/data crawler. Deprecated Anthropic crawler token. Harmless to leave blocked, but it is no longer the token Anthropic documents.
- [Google-Extended](https://checkaibots.com/crawler/google-extended): Google's training/data crawler. Not a crawler — a robots.txt control token with no separate HTTP user-agent. Controls whether crawled content may be used for Gemini training and Gemini grounding. Google states it does not affect Search inclusion or ranking.
- [GoogleOther](https://checkaibots.com/crawler/googleother): Google's training/data crawler. Google's generic crawler used for one-off research and internal R&D fetches. Google states it does not affect any specific product.
- [Applebot-Extended](https://checkaibots.com/crawler/applebot-extended): Apple's training/data crawler. Not a crawler — a control token that decides how Apple may use content already crawled by Applebot. Opts you out of Apple foundation-model training.
- [CCBot](https://checkaibots.com/crawler/ccbot): Common Crawl's training/data crawler. Feeds the Common Crawl open dataset, which many model trainers build on.
- [Bytespider](https://checkaibots.com/crawler/bytespider): ByteDance's training/data crawler. ByteDance data collection. Frequently reported to ignore robots.txt — community-reported; ByteDance publishes limited crawler documentation.
- [Meta-ExternalAgent](https://checkaibots.com/crawler/meta-externalagent): Meta's training/data crawler. Meta AI training and data collection.
- [FacebookBot](https://checkaibots.com/crawler/facebookbot): Meta's training/data crawler. Meta's older crawler token, largely superseded by Meta-ExternalAgent.
- [Amazonbot](https://checkaibots.com/crawler/amazonbot): Amazon's training/data crawler. Amazon describes it as powering Alexa answers; grouped here as data collection, though Amazon's own wording is closer to retrieval.
- [cohere-ai](https://checkaibots.com/crawler/cohere-ai): Cohere's training/data crawler. Cohere model training crawler.
- [cohere-training-data-crawler](https://checkaibots.com/crawler/cohere-training-data-crawler): Cohere's training/data crawler. Cohere legacy training-data crawler token.
- [Diffbot](https://checkaibots.com/crawler/diffbot): Diffbot's training/data crawler. Commercial knowledge-graph and data-broker crawler; blocking prevents content resale.
- [MistralAI-Training](https://checkaibots.com/crawler/mistralai-training): Mistral's training/data crawler. Mistral model training crawler.
- [YouBot](https://checkaibots.com/crawler/youbot): You.com's training/data crawler. You.com index and training crawler.
- [Webzio-Extended](https://checkaibots.com/crawler/webzio-extended): Webz.io's training/data crawler. Commercial data broker licensing content for AI training.
- [AI2Bot](https://checkaibots.com/crawler/ai2bot): Allen Institute for AI's training/data crawler. Academic research crawler (Allen AI); builds open datasets such as Dolma.
- [PanguBot](https://checkaibots.com/crawler/pangubot): Huawei's training/data crawler. Huawei Pangu model training.
- [Timpibot](https://checkaibots.com/crawler/timpibot): Timpi's training/data crawler. Timpi independent web index (search and AI data).
- [DeepSeek-Bot](https://checkaibots.com/crawler/deepseek-bot): DeepSeek's training/data crawler. DeepSeek model training and data collection.
- [OAI-SearchBot](https://checkaibots.com/crawler/oai-searchbot): OpenAI's retrieval/search crawler. OpenAI's search crawler. Sites opted out of it are not shown in ChatGPT search answers — allow it if you want citations.
- [OAI-AdsBot](https://checkaibots.com/crawler/oai-adsbot): OpenAI's retrieval/search crawler. Visits landing pages submitted as ChatGPT ads for policy review. Only those pages, and the data is not used to train foundation models.
- [Claude-SearchBot](https://checkaibots.com/crawler/claude-searchbot): Anthropic's retrieval/search crawler. Anthropic's search crawler; disabling it may reduce your visibility and accuracy in Claude search results.
- [PerplexityBot](https://checkaibots.com/crawler/perplexitybot): Perplexity's retrieval/search crawler. Perplexity's search crawler. Perplexity recommends allowing it plus its published IP ranges; it is not used to train foundation models.
- [Googlebot](https://checkaibots.com/crawler/googlebot): Google's retrieval/search crawler. Google's main search crawler; controls Google Search, including AI features built on the Search index. Almost never block.
- [Bingbot](https://checkaibots.com/crawler/bingbot): Microsoft's retrieval/search crawler. Bing index and Microsoft Copilot retrieval; almost never block.
- [Applebot](https://checkaibots.com/crawler/applebot): Apple's retrieval/search crawler. Powers Spotlight, Siri and Safari search. Training use is governed by the separate Applebot-Extended token.
- [DuckAssistBot](https://checkaibots.com/crawler/duckassistbot): DuckDuckGo's retrieval/search crawler. DuckDuckGo AI summaries.
- [meta-webindexer](https://checkaibots.com/crawler/meta-webindexer): Meta's retrieval/search crawler. Meta AI search index; blocking removes you from Meta AI answers.
- [PetalBot](https://checkaibots.com/crawler/petalbot): Huawei's retrieval/search crawler. Petal Search / Celia assistant index.
- [ImagesiftBot](https://checkaibots.com/crawler/imagesiftbot): Hive's retrieval/search crawler. Hive's AI image search index.
- [iaskspider](https://checkaibots.com/crawler/iaskspider): iAsk.ai's retrieval/search crawler. iAsk.ai AI search engine index.
- [ChatGPT-User](https://checkaibots.com/crawler/chatgpt-user): OpenAI's user-triggered fetcher. Fetches a page when someone opens a link inside ChatGPT. OpenAI states robots.txt rules may not apply because the request is user-triggered.
- [Claude-User](https://checkaibots.com/crawler/claude-user): Anthropic's user-triggered fetcher. User-initiated fetch for Claude; disabling it may reduce visibility in user-directed search.
- [Perplexity-User](https://checkaibots.com/crawler/perplexity-user): Perplexity's user-triggered fetcher. User-initiated fetch. Perplexity states this fetcher generally ignores robots.txt rules; it is not a crawler and is not used for training.
- [MistralAI-User](https://checkaibots.com/crawler/mistralai-user): Mistral's user-triggered fetcher. User-initiated fetch when Le Chat users open links.
- [Meta-ExternalFetcher](https://checkaibots.com/crawler/meta-externalfetcher): Meta's user-triggered fetcher. User-initiated fetch when Meta AI users open links.
- [DeepSeek-User](https://checkaibots.com/crawler/deepseek-user): DeepSeek's user-triggered fetcher. User-initiated fetch when DeepSeek users open links.

## Guides

- [Cloudflare Blocks AI Crawlers by Default on September 15 — What It Means for Your Site](https://checkaibots.com/blog/cloudflare-september-15-ai-crawler-block): Starting September 15, 2026, Cloudflare begins blocking AI training and agent crawlers by default on new domains — even if your robots.txt says allow. Here's what changes, what it means for your AI citations, and how to check which side your site is on.
- [What Is OAI-SearchBot? The Crawler That Gets You Cited in ChatGPT](https://checkaibots.com/blog/what-is-oai-searchbot): OAI-SearchBot is OpenAI's retrieval crawler — the one that decides whether your site gets cited in ChatGPT answers. Blocking it costs you citations. Here's what it is and how to check yours.
- [GPTBot Blocked? Here's What to Do (and What NOT to Do)](https://checkaibots.com/blog/gptbot-blocked-fix): If your robots.txt blocks GPTBot, you've opted out of OpenAI model training — which may be intentional. The real danger is accidentally blocking the wrong crawler along the way. Here's how to fix it the right way.
- [Is My Site Visible to ChatGPT? A 5-Minute Audit](https://checkaibots.com/blog/is-my-site-visible-to-chatgpt): Three fast checks to see whether ChatGPT can access and cite your site — and what to fix if it can't. No tools needed for the first two.
- [How to Detect AI Crawlers Visiting Your Website](https://checkaibots.com/blog/detect-ai-crawlers-visiting-website): Want to know if GPTBot, ClaudeBot or PerplexityBot are visiting your site? Three practical methods: server logs, robots.txt checks, and a crawler visibility scan.
- [AI Crawlers 101: What's Crawling Your Site in 2026](https://checkaibots.com/blog/ai-crawlers-101): GPTBot, ClaudeBot, PerplexityBot, OAI-SearchBot... there are now dozens of AI crawlers, and they do very different jobs. Here's what each type does and why it matters for your site.
- [The Most Expensive Mistake in AI SEO: Blocking the Wrong Crawler](https://checkaibots.com/blog/expensive-mistake-ai-seo): Trying to block GPTBot (training) but accidentally blocking OAI-SearchBot (retrieval)? That's the mistake that quietly removes you from ChatGPT answers while doing nothing to protect your content.
- [How to Check If ChatGPT Can See Your Site](https://checkaibots.com/blog/check-if-chatgpt-can-see-your-site): Three ways to check whether your site is visible to AI search: the quick robots.txt check, the UA-level fetch check, and what to do when you find a problem.
