GPTBot — OpenAI training crawler
A technical reference for identifying and governing GPTBot access.
AI Summary: GPTBot is treated as a training crawler in the catalog. A
Disallowrule is evidence of your declared preference, not a guarantee about vendor behavior.
Role and policy boundary
GPTBot represents a training-oriented crawler in this catalog. That distinction matters because a site may want to remain discoverable in AI search while declining the use of content for model training. Do not rely on a wildcard policy if those intents differ. Give training agents their own blocks, document the decision, and review the rendered response headers as a second evidence layer.
The minimum robots policy to block this crawler is:
User-agent: GPTBot
Disallow: /
To allow the crawler across the public site while keeping a private area closed, use an explicit allow followed by a scoped disallow:
User-agent: GPTBot
Allow: /
Disallow: /private/
Disallow: /admin/
The policy engine treats the path match and the exact user-agent as evidence. A missing GPTBot group does not become an automatic allow; it is reported as unknown when the remaining evidence is incomplete.
Layered verification
robots.txt is only one layer. A page can still send a restrictive meta directive or response header:
<meta name="robots" content="noindex, nofollow">
X-Robots-Tag: noindex, nofollow
In an AI Bot Check report, a GPTBot Allow in robots.txt combined with noindex in a response header is a cross-layer conflict. The report keeps the raw line, normalized directive, source layer, and affected path together so an operator can fix the correct system instead of editing a random file.
WAF and Nginx remediation examples
A Cloudflare Custom Rule can start as an advisory rule that logs the crawler before enforcement. Replace the action only after observing legitimate traffic:
{
"description": "Review GPTBot training access",
"expression": "lower(http.user_agent) contains \"gptbot\"",
"action": "log",
"note": "Change to block only after validating the business policy."
}
For Nginx, a map makes the decision visible and composable with an existing server policy:
map $http_user_agent $gptbot_action {
default allow;
~*"GPTBot" block;
}
These snippets are starting points, not vendor guarantees. Test them with a controlled request and verify that caches, CDNs, and origin headers do not reintroduce contradictory signals.
Review checklist
Check the exact casing-independent user-agent, confirm that the intended path is covered, inspect meta and X-Robots-Tag, and record the policy change with a date. If the site has separate search and training goals, review GPTBot beside OAI-SearchBot rather than collapsing both into one wildcard rule.
Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.