Bot directory / AI Training Crawler

GPTBot — OpenAI training crawler

A technical reference for identifying and governing GPTBot access.

AI Summary: GPTBot is treated as a training crawler in the catalog. A Disallow rule is evidence of your declared preference, not a guarantee about vendor behavior.

Role and policy boundary

GPTBot represents a training-oriented crawler in this catalog. That distinction matters because a site may want to remain discoverable in AI search while declining the use of content for model training. Do not rely on a wildcard policy if those intents differ. Give training agents their own blocks, document the decision, and review the rendered response headers as a second evidence layer.

The minimum robots policy to block this crawler is:

configuration / code
User-agent: GPTBot
Disallow: /

To allow the crawler across the public site while keeping a private area closed, use an explicit allow followed by a scoped disallow:

configuration / code
User-agent: GPTBot
Allow: /
Disallow: /private/
Disallow: /admin/

The policy engine treats the path match and the exact user-agent as evidence. A missing GPTBot group does not become an automatic allow; it is reported as unknown when the remaining evidence is incomplete.

Layered verification

robots.txt is only one layer. A page can still send a restrictive meta directive or response header:

configuration / code
<meta name="robots" content="noindex, nofollow">
configuration / code
X-Robots-Tag: noindex, nofollow

In an AI Bot Check report, a GPTBot Allow in robots.txt combined with noindex in a response header is a cross-layer conflict. The report keeps the raw line, normalized directive, source layer, and affected path together so an operator can fix the correct system instead of editing a random file.

WAF and Nginx remediation examples

A Cloudflare Custom Rule can start as an advisory rule that logs the crawler before enforcement. Replace the action only after observing legitimate traffic:

configuration / code
{
  "description": "Review GPTBot training access",
  "expression": "lower(http.user_agent) contains \"gptbot\"",
  "action": "log",
  "note": "Change to block only after validating the business policy."
}

For Nginx, a map makes the decision visible and composable with an existing server policy:

configuration / code
map $http_user_agent $gptbot_action {
    default allow;
    ~*"GPTBot" block;
}

These snippets are starting points, not vendor guarantees. Test them with a controlled request and verify that caches, CDNs, and origin headers do not reintroduce contradictory signals.

Review checklist

Check the exact casing-independent user-agent, confirm that the intended path is covered, inspect meta and X-Robots-Tag, and record the policy change with a date. If the site has separate search and training goals, review GPTBot beside OAI-SearchBot rather than collapsing both into one wildcard rule.


Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.