← Bot Directory/OAI-SearchBot
Bot directory / AI Search Bot

OAI-SearchBot — OpenAI search crawler

A technical reference for OpenAI search discovery policy.

AI Summary: OAI-SearchBot belongs to the search group. Govern it separately from GPTBot so search visibility and training preferences do not become one accidental rule.

Role and policy boundary

Treat this User-Agent as a distinct policy identity. The correct decision depends on the bot purpose, the paths it requests, and the response controls applied by the origin, CDN, and application layer. A robots rule is a declaration of intent; it does not replace authentication, authorization, or rate limiting.

OAI-SearchBot is catalogued as a search-oriented crawler. A search crawler can be allowed even when a training crawler is blocked, but that outcome is only explainable when the rules are explicit. Start with a dedicated group:

configuration / code
User-agent: OAI-SearchBot
Allow: /
Disallow: /staging/
Disallow: /internal/

To stop search discovery for the entire site, use:

configuration / code
User-agent: OAI-SearchBot
Disallow: /

Avoid assuming that User-agent: * expresses the same business intent. A wildcard can affect assistant and training crawlers too, and it makes later audits harder because the source of the decision is less specific.

Layered verification

Verify the same URL through each control plane instead of assuming that one green signal represents the whole request path. Compare the bot-specific robots group, the page-level metadata, and the response headers captured at the public edge.

A common policy is to keep search access open while refusing training access:

configuration / code
User-agent: OAI-SearchBot
Allow: /

User-agent: GPTBot
Disallow: /

The Policy Engine evaluates the selected user-agent, path scope, and the other supplied layers independently. It can therefore explain why OAI-SearchBot is allowed while GPTBot is blocked, rather than returning one blended website score.

Page-level directives can still override the intended outcome for indexing:

configuration / code
<meta name="robots" content="index, follow">
configuration / code
X-Robots-Tag: index, follow

If a response uses noindex, the report marks the result as blocked or conflicting even if the crawler-specific robots group is permissive. This is especially important for canonical pages served through an edge cache where headers may differ from the origin response.

WAF and Nginx remediation examples

A Cloudflare rule can log OAI-SearchBot for observability without blocking it:

configuration / code
{
  "description": "Observe OAI-SearchBot discovery",
  "expression": "lower(http.user_agent) contains \"oai-searchbot\"",
  "action": "log",
  "note": "Keep search access open while reviewing request quality."
}

A Vercel middleware seam can attach a diagnostic header while leaving the request available to the application:

configuration / code
export function middleware(request: Request) {
  const userAgent = request.headers.get('user-agent')?.toLowerCase() ?? '';
  const response = new Response(null, { status: 204 });
  if (userAgent.includes('oai-searchbot')) response.headers.set('x-aibotcheck-policy', 'search-allow');
  return response;
}

Use your platform’s actual middleware response pattern rather than copying this simplified example without review. Never place a secret, verification token, or internal policy identifier in a public response header.

configuration / code
map $http_user_agent $block_oai_searchbot_private {
    default 0;
    ~*oai-searchbot 1;
}

server {
    location ~ ^/(admin|account|private|licensed|internal|api)/ {
        if ($block_oai_searchbot_private) { return 403; }
        try_files $uri $uri/ =404;
    }
}

Review checklist

Use this checklist after every policy change and after a CDN or WAF migration. Record the request URL, User-Agent, HTTP status, final redirect, and the exact evidence used to reach the decision.

Verify that the dedicated group appears before relying on a wildcard, test a representative public and private path, and compare live response headers with robots.txt. Keep the search policy close to the content owner’s intent and record whether the site wants discovery, citation, or no access at all.


Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.