Open dataset feeding most open-source model training
Your call — blocking is a legitimate choice
CCBot is Common Crawl's training/data crawler. Blocking it opts your content out of Common Crawl's model training and datasets. It does not remove you from AI answers — citations are governed by separate retrieval crawlers.
Common reasons to block: copyright or licensing policy, competitive sensitivity, paywalled content.
Common reasons to allow: maximum AI reach, research visibility, and being part of the datasets future models learn from.
User-agent: CCBot Allow: /
User-agent: CCBot Disallow: /
See whether CCBot — and 39 other AI crawlers — can reach your site today.