Bot directory / ai-training

How to block CCBot: The Common Crawl AI Training Crawler

A complete technical reference for CCBot, the web crawler operated by Common Crawl to build datasets used for training large language models.

AI Summary: CCBot is the primary web crawler for Common Crawl, an open repository of web crawl data widely used to train AI models like GPT and Claude. To block CCBot from using your site's data for AI training, add User-agent: CCBot followed by Disallow: / to your robots.txt file, or use firewall rules to block its User-Agent string.

Role and policy boundary

CCBot is operated by Common Crawl, a non-profit organization that builds and maintains an open repository of web crawl data. Unlike search engine crawlers that index content to drive traffic back to your site, CCBot's primary purpose is data collection for large-scale datasets.

These datasets, such as the Common Crawl corpus, are foundational for training Large Language Models (LLMs) across the AI industry. Because Common Crawl data is open and distributed, allowing CCBot means your content may be ingested by multiple AI vendors simultaneously. Blocking CCBot is a critical step in a comprehensive AI training opt-out strategy, as it cuts off one of the largest upstream data sources for model training.

Layered verification

When implementing a block for CCBot, it is essential to use a layered approach to ensure the crawler respects your directives across all access points.

The standard and most recognized method is using the robots.txt file. Common Crawl explicitly states that CCBot respects standard robots.txt directives. However, because CCBot fetches data at scale, relying solely on robots.txt may leave your site vulnerable if the file is temporarily unavailable or misconfigured.

To reinforce this, you should also implement HTTP headers or meta tags. The X-Robots-Tag: noai, noimageai header provides a strong signal that the content should not be used for AI training, though its support varies among downstream consumers of the Common Crawl dataset. For absolute certainty, blocking the User-Agent at the network edge ensures the crawler cannot access the content at all.

WAF and Nginx remediation examples

To block CCBot effectively at the server or network level, you can implement rules that reject requests matching its User-Agent string.

For Nginx servers, you can add a conditional block in your configuration file to return a 403 Forbidden status when CCBot is detected:

configuration / code
if ($http_user_agent ~* (CCBot)) {
    return 403;
}

If you are using Cloudflare WAF, you can deploy a custom firewall rule using the following expression to block the crawler before it reaches your origin server:

configuration / code
{
  "action": "block",
  "expression": "(http.user_agent contains \"CCBot\")",
  "description": "Block Common Crawl CCBot for AI training opt-out"
}

Review checklist

Before finalizing your policy for CCBot, review the following checkpoints to ensure comprehensive coverage:

  1. Verify that your robots.txt includes a specific User-agent: CCBot block.
  2. Confirm that edge firewall rules (WAF) are active and correctly identifying the CCBot User-Agent string.
  3. Test your configuration by simulating a request using the CCBot User-Agent to ensure it returns a 403 or 401 status code.
  4. Ensure that blocking CCBot aligns with your broader data governance and AI training opt-out policies.

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.