Training Opt-Out Policy: How to Block AI Data Scraping
A comprehensive guide to blocking AI crawlers from using your website's content for training large language models (LLMs) while maintaining search visibility.
AI Summary: A training opt-out policy allows website owners to declare that their content must not be used to train AI models. Because AI companies use different crawlers for search indexing and model training, webmasters can use
robots.txtto block training bots (likeGPTBotorApplebot-Extended) while still allowing search bots (likeOAI-SearchBotorApplebot) to index their site for visibility.
The Separation of Search and Training
In the era of generative AI, content discovery and content extraction for model training are two distinct operations. Many major AI vendors now operate separate crawlers or provide specific robots.txt tokens to distinguish between these intents.
If you block a vendor's primary search crawler, your website may disappear from their search engine or AI-powered answer engine entirely. If you only block their training crawler, your site remains visible in search results, but the vendor is instructed not to scrape your content to train future models.
For example:
- OpenAI uses
OAI-SearchBotfor SearchGPT visibility, butGPTBotfor model training. - Apple uses
Applebotfor Siri and Spotlight search, butApplebot-Extendedas an opt-out token for generative AI training. - Anthropic uses
ClaudeBotfor general web fetching and training.
Recommended Layering for Training Opt-Outs
To effectively opt out of AI training without sacrificing search visibility, you must explicitly declare rules for known training crawlers in your robots.txt file.
Step 1: Identify Training Crawlers
Maintain a list of known AI training crawlers. Common examples include:
GPTBot(OpenAI)Applebot-Extended(Apple)ClaudeBot(Anthropic)Bytespider(ByteDance)CCBot(Common Crawl)FacebookBot(Meta)Google-Extended(Google's token for Bard/Vertex AI training)
Step 2: Configure robots.txt
Add explicit Disallow rules for these crawlers. Ensure these rules are placed before any wildcard (*) rules.
# Block OpenAI's training crawler
User-agent: GPTBot
Disallow: /
# Block Apple's AI training token
User-agent: Applebot-Extended
Disallow: /
# Block Google's AI training token
User-agent: Google-Extended
Disallow: /
# Block Anthropic's crawler
User-agent: ClaudeBot
Disallow: /
# Block Common Crawl (often used by open-source models)
User-agent: CCBot
Disallow: /
# Allow general search engines
User-agent: *
Allow: /
Enforcement and WAF Remediation
While robots.txt is the standard method for declaring a training opt-out, it is purely advisory. Malicious or poorly-configured scrapers may ignore these directives.
To enforce your policy, you must implement network-level controls using a Web Application Firewall (WAF) or server configuration (like Nginx).
WAF JSON Example
You can create a WAF rule to block requests from known training User-Agents:
{
"description": "Block known AI training crawlers",
"expression": "lower(http.user_agent) contains \"gptbot\" or lower(http.user_agent) contains \"claudebot\" or lower(http.user_agent) contains \"bytespider\" or lower(http.user_agent) contains \"ccbot\"",
"action": "block"
}
(Note: Applebot-Extended and Google-Extended are policy tokens, not User-Agents, so they will not appear in HTTP headers and cannot be blocked via WAF User-Agent rules. You must block the primary crawler, e.g., Applebot or Googlebot, if you wish to enforce the block at the network level, though this will also block search indexing.)
Nginx Remediation Example
You can block these agents directly in your web server configuration:
map $http_user_agent $block_ai_training {
default 0;
~*GPTBot 1;
~*ClaudeBot 1;
~*Bytespider 1;
~*CCBot 1;
~*FacebookBot 1;
}
server {
location / {
if ($block_ai_training) {
return 403;
}
# Standard configuration follows...
}
}
Review Checklist
When implementing a training opt-out policy, verify the following:
- [ ] Have you separated search crawlers (e.g.,
OAI-SearchBot) from training crawlers (e.g.,GPTBot) in yourrobots.txt? - [ ] Are the training crawlers explicitly disallowed (
Disallow: /)? - [ ] Have you implemented WAF or server-level blocks for crawlers that are known to ignore
robots.txtor for which you require strict enforcement? - [ ] Have you tested your site's visibility on major search engines to ensure you haven't accidentally blocked general indexing?
- [ ] Are you monitoring your server logs for new or spoofed AI crawler User-Agents?
Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.