Identifying Spoofed AI Bots: Verification & Forensics
A technical forensics guide to detecting and blocking malicious scrapers that forge User-Agent headers to masquerade as legitimate AI crawlers.
AI Summary: Publishers detect spoofed AI bots by combining Forward-Confirmed Reverse DNS (FCrDNS), official ASN verification, and TLS fingerprinting (JA4). User-Agent headers alone cannot be trusted; any request failing cryptographic DNS verification must be blocked at the edge.
The Reality of User-Agent Spoofing
The HTTP User-Agent header is client-controlled text. Anyone can send a request with:
User-Agent: Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)
Malicious scrapers, competitive intelligence harvesters, and vulnerability scanners routinely impersonate legitimate search bots and AI assistants to evade paywalls, bypass rate limits, and prevent IP bans.
Relying solely on user-agent strings creates severe security vulnerabilities and skews bot analytics.
The Triad of Crawler Verification
Legitimate AI crawler identification requires three independent verification layers:
1. User-Agent String Check (Initial Filter)
│
▼
2. Forward-Confirmed Reverse DNS (FCrDNS) (Cryptographic Validation)
│
▼
3. Autonomous System Number (ASN) & Official IP Prefix Check (Network Layer)
Method 1: Forward-Confirmed Reverse DNS (FCrDNS)
This two-step DNS lookup verifies domain authenticity:
# Step 1: Perform Reverse DNS (PTR) lookup on client IP
PTR=$(dig -x 20.171.207.1 +short | sed 's/\.$//')
echo "PTR hostname: $PTR"
# Must match vendor domain pattern: *.openai.com, *.googlebot.com, *.search.msn.com
# Step 2: Perform Forward DNS (A) lookup on the returned hostname
RESOLVED_IP=$(dig +short "$PTR")
echo "Forward resolved IP: $RESOLVED_IP"
# Step 3: Compare original IP with forward resolved IP
if [ "20.171.207.1" = "$RESOLVED_IP" ]; then
echo "VERIFIED: Genuine crawler identity confirmed."
else
echo "ALERT: SPOOFED BOT DETECTED. Reject immediately."
fi
Method 2: Official IP Prefix Verification
If DNS queries introduce unacceptable latency at your edge proxy, validate the client IP against vendor-published IP range JSON feeds:
| Crawler | Official IP Feed URL | Expected Domain Suffix |
| :--- | :--- | :--- |
| OpenAI | https://openai.com/gptbot.json | openai.com |
| Google | https://developers.google.com/search/apis/ipranges/googlebot.json | googlebot.com |
| Bing | Microsoft Download Center JSON feeds | search.msn.com |
Automated WAF Action for Spoofed Bots
Configure edge WAF logic to block requests where the User-Agent claims to be an AI bot but fails reverse DNS:
{
"description": "Block fake OpenAI crawler user-agent from unauthorized network",
"action": "block",
"expression": "(http.user_agent contains "GPTBot" and not ip.src in {20.171.207.0/24 20.171.206.0/24})"
}
Protect your origin servers from malicious scrapers masquerading as AI crawlers. Audit your bot security with Geolify.ai.