Bot directory / search-engine

Neevabot: Robots.txt & Crawl Policy Reference

Technical reference for the historical Neevabot registry label, with explicit limits around the unresolved neeva.com source and current crawler policy.

AI Summary: Neevabot/1.0 is a historical registry label associated with Neeva search crawling, but the linked neeva.com/neevabot source failed DNS resolution during review and a headless Bing search returned unrelated results. No current policy, robots behavior, IP range, rate guidance, or verification method was established; treat the header as an unverified historical signal.

Role and policy boundary

The registry describes Neevabot as a Neeva search-engine web crawler and records this User-Agent:

configuration / code
Mozilla/5.0 (compatible; Neevabot/1.0; +https://neeva.com/neevabot)

The registry-linked https://neeva.com/neevabot page was attempted with the headless browser and failed with net::ERR_NAME_NOT_RESOLVED; DNS availability could not be confirmed from this environment. A follow-up headless Bing search for Neevabot and Neeva shutdown/policy terms returned unrelated insurance results and no verified Neevabot source in the extracted content.

The Neeva search-crawling role is therefore a historical registry description only. A request carrying this token may come from an old deployment, a third-party tool, a test harness, a fork, or a spoofed header. Do not infer current activity, operator identity, downstream use, AI-training purpose, permission to crawl, or access to private content from the label, the URL in the header, DNS failure, or unrelated search results.

If logs later establish the exact client and your policy is to exclude it, a narrow defensive group could be:

configuration / code
User-agent: Neevabot
Disallow: /

For selective access after confirmation:

configuration / code
User-agent: Neevabot
Allow: /public/
Allow: /docs/
Disallow: /admin/
Disallow: /account/
Disallow: /private/
Disallow: /api/

These are site-owner examples, not recovered Neeva instructions. Robots.txt is advisory and cannot protect private or licensed content; use authentication, authorization, signed URLs, data minimization, and origin controls for those boundaries.

Layered verification

Begin with raw access logs and preserve the complete User-Agent, source IP, ASN, reverse-DNS result, HTTP method, requested path, response status, response size, redirect chain, timestamp, and request rate. No accessible official source in this review published a current Neevabot IP range, reverse-DNS procedure, Crawl-delay, rate policy, or opt-out process. The registered header is useful for triage but is not authentication.

Do not treat the hostname inside the User-Agent as proof of ownership. Confirm any future source independently, and do not use a current search-crawler verification procedure for Neevabot unless it explicitly names this token.

Compare observed behavior with a search-indexing hypothesis without turning it into attribution. Public HTML, canonical metadata, feeds, and sitemaps may be consistent with search discovery. Private endpoints, account routes, licensed documents, APIs, high concurrency, repeated retries, and unexpected downloads may indicate spoofing, abuse, a partner integration, or another client. These observations establish impact and data risk, not operator identity or downstream use.

Evaluate /robots.txt only after the actual token is known. Confirm it is served from the intended host, returns a successful status and text content type, and contains the exact group you intend to publish. Test the complete registered header, any confirmed Neevabot token, and the global group separately. Page-level directives can express discovery preferences:

configuration / code
<meta name="robots" content="noindex, nofollow">
configuration / code
X-Robots-Tag: noindex, nofollow

These signals do not authenticate Neevabot or secure private routes. Enforce sensitive boundaries in the application and at the origin.

WAF and Nginx remediation examples

When the identity is unknown, use report-only logging and avoid a broad neeva or bot substring rule. Adapt the expression to your WAF provider:

configuration / code
{
  "description": "Log historical Neevabot candidates for verification",
  "expression": "lower(http.user_agent) contains \"neevabot\"",
  "action": "log"
}

Once a complete token and authoritative source establish the identity, use a narrow route-scoped control:

configuration / code
map $http_user_agent $block_confirmed_neevabot_private {
    default 0;
    # Add only a complete, independently verified token here.
    # ~*Confirmed-Neevabot-Token 1;
}

server {
    location ~ ^/(admin|account|private|licensed|internal|api)/ {
        if ($block_confirmed_neevabot_private) { return 403; }
        try_files $uri $uri/ =404;
    }
}

A User-Agent match is easy to spoof and may catch an authorized integration or a local test. Do not invent an IP allowlist, reverse-DNS suffix, or vendor exception from the token or DNS error. Test public pages, feeds, sitemaps, accounts, APIs, licensed documents, and approved integrations separately. Pair edge controls with authentication, rate limits, signed assets, caching, data-loss monitoring, and anomaly detection.

Review checklist

Search logs for the complete registered header and preserve representative source IPs, ASNs, reverse-DNS results, paths, methods, response sizes, statuses, timing, and rate. Check whether a current operator source can independently verify the client before assigning Neeva ownership or search purpose.

Re-check neeva.com, the /neevabot path, the community registry, and current search-crawler documentation when a credible source becomes reachable. Keep this profile at legacy-label until a first-party source publishes a current token, purpose, source-verification method, rate guidance, robots behavior, or opt-out process. Treat DNS failure and unrelated Bing results as source-discovery limitations, not proof of inactivity or shutdown.

Decide whether your objective is to preserve public discovery, limit extraction, protect private material, or reduce crawl load. Publish exact robots rules only after the token is known, and enforce private routes with application and WAF controls. Do not claim successful blocking from configuration alone; verify subsequent logs and response behavior.

The registry’s search-indexing label does not establish AI-model training use. Keep any downstream-use statement separate from the limited evidence available here.

References

  1. Neevabot documentation — registry-linked source; headless-browser request failed with net::ERR_NAME_NOT_RESOLVED on 2026-08-25.
  2. Headless Bing search for Neevabot — follow-up search returned unrelated results and no verified Neevabot source in the extracted content on 2026-08-25.
  3. Crawler User Agents community registry — registry context for the label and header; it does not authenticate current Neeva infrastructure.
  4. Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
  5. RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.