Nginx Configuration: Blocking & Throttling AI Bots
High-performance Nginx map directives, rate limits, and status codes to block, rate-limit, and segregate AI crawlers at the edge.
AI Summary: Nginx handles AI bot governance with minimal CPU overhead using the map directive inside the http block. By mapping the $http_user_agent variable to binary flags, Nginx rejects unwanted crawlers with HTTP 403 or 429 before proxying requests upstream.
High-Performance Bot Governance in Nginx
Nginx is the de facto edge reverse proxy for modern web applications. Managing automated crawler traffic in Nginx prevents backend application pools (such as Node.js, Gunicorn, or PHP-FPM) from exhausting thread pools during aggressive bot sweeps.
Using Nginx's map directive compiles user-agent strings into a hash table at startup, delivering $O(1)$ lookup performance on every incoming HTTP request.
Step 1: Define the Crawler Map in the http Block
Add this block inside /etc/nginx/nginx.conf (within the http { ... } block):
# Map client User-Agent against known AI training crawlers
map $http_user_agent $blocked_ai_crawler {
default 0;
# Major AI Training Bots
~*GPTBot 1;
~*ClaudeBot 1;
~*Bytespider 1;
~*CCBot 1;
~*Amazonbot 1;
~*Diffbot 1;
~*ImagesiftBot 1;
~*cohere-ai 1;
}
Step 2: Enforce Blocking in the Server Block
In your virtual host configuration file (e.g., /etc/nginx/sites-available/default):
server {
listen 443 ssl http2;
server_name example.com;
# Immediate rejection of mapped crawlers
if ($blocked_ai_crawler) {
return 403 "Access denied: AI training crawler blocked by server policy.";
}
location / {
proxy_pass http://backend_upstream;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
}
}
Step 3: Throttling Search Crawlers with Rate Limiting
If you want to permit AI search crawlers (like OAI-SearchBot or PerplexityBot) but prevent aggressive spikes from degrading user performance, apply dedicated rate limiting zones:
# Define rate limiting zone for AI search crawlers
limit_req_zone $binary_remote_addr zone=ai_search_zone:10m rate=5r/s;
server {
location / {
# Check if client is an AI search bot
if ($http_user_agent ~* (OAI-SearchBot|PerplexityBot)) {
limit_req zone=ai_search_zone burst=10 nodelay;
}
proxy_pass http://backend_upstream;
}
}
Verification & Syntax Testing
# 1. Test Nginx syntax
nginx -t
# 2. Reload Nginx without dropping active connections
systemctl reload nginx
# 3. Validate enforcement with curl
curl -I -A "Mozilla/5.0 (compatible; ClaudeBot/1.0; +https://www.anthropic.com/claudebot)" https://example.com/
# Output must be: HTTP/2 403
Protect your Nginx infrastructure from aggressive crawler exhaustion. Analyze your bot posture with Geolify.ai.