← Policy library/Identifying Spoofed AI Bots: Verification & Forensics
Policy library

Identifying Spoofed AI Bots: Verification & Forensics

A technical forensics guide to detecting and blocking malicious scrapers that forge User-Agent headers to masquerade as legitimate AI crawlers.

AI Summary: Publishers detect spoofed AI bots by combining Forward-Confirmed Reverse DNS (FCrDNS), official ASN verification, and TLS fingerprinting (JA4). User-Agent headers alone cannot be trusted; any request failing cryptographic DNS verification must be blocked at the edge.

The Reality of User-Agent Spoofing

The HTTP User-Agent header is client-controlled text. Anyone can send a request with:

configuration / code
User-Agent: Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)

Malicious scrapers, competitive intelligence harvesters, and vulnerability scanners routinely impersonate legitimate search bots and AI assistants to evade paywalls, bypass rate limits, and prevent IP bans.

Relying solely on user-agent strings creates severe security vulnerabilities and skews bot analytics.

The Triad of Crawler Verification

Legitimate AI crawler identification requires three independent verification layers:

configuration / code
1. User-Agent String Check (Initial Filter)
                 │
                 ▼
2. Forward-Confirmed Reverse DNS (FCrDNS) (Cryptographic Validation)
                 │
                 ▼
3. Autonomous System Number (ASN) & Official IP Prefix Check (Network Layer)

Method 1: Forward-Confirmed Reverse DNS (FCrDNS)

This two-step DNS lookup verifies domain authenticity:

configuration / code
# Step 1: Perform Reverse DNS (PTR) lookup on client IP
PTR=$(dig -x 20.171.207.1 +short | sed 's/\.$//')
echo "PTR hostname: $PTR"
# Must match vendor domain pattern: *.openai.com, *.googlebot.com, *.search.msn.com

# Step 2: Perform Forward DNS (A) lookup on the returned hostname
RESOLVED_IP=$(dig +short "$PTR")
echo "Forward resolved IP: $RESOLVED_IP"

# Step 3: Compare original IP with forward resolved IP
if [ "20.171.207.1" = "$RESOLVED_IP" ]; then
    echo "VERIFIED: Genuine crawler identity confirmed."
else
    echo "ALERT: SPOOFED BOT DETECTED. Reject immediately."
fi

Method 2: Official IP Prefix Verification

If DNS queries introduce unacceptable latency at your edge proxy, validate the client IP against vendor-published IP range JSON feeds:

| Crawler | Official IP Feed URL | Expected Domain Suffix | | :--- | :--- | :--- | | OpenAI | https://openai.com/gptbot.json | openai.com | | Google | https://developers.google.com/search/apis/ipranges/googlebot.json | googlebot.com | | Bing | Microsoft Download Center JSON feeds | search.msn.com |

Automated WAF Action for Spoofed Bots

Configure edge WAF logic to block requests where the User-Agent claims to be an AI bot but fails reverse DNS:

configuration / code
{
  "description": "Block fake OpenAI crawler user-agent from unauthorized network",
  "action": "block",
  "expression": "(http.user_agent contains "GPTBot" and not ip.src in {20.171.207.0/24 20.171.206.0/24})"
}

Protect your origin servers from malicious scrapers masquerading as AI crawlers. Audit your bot security with Geolify.ai.

Related policies