Bot directory / search-engine

OICrawler: Robots.txt & Crawl Policy Reference

Technical reference for the OICrawler registry label and current OpenIndex discovery service, with explicit limits around User-Agent and crawler verification.

AI Summary: The registry records OICrawler/Nutch https://openindex.ai and describes it as an OpenIndex indexing crawler, but the current OpenIndex site redirects to chat.openindex.ai and documents encrypted messaging and AI-agent discovery rather than a crawler specification. Its current robots file has only a global group blocking /login and /logout; it has no OICrawler group, Crawl-delay, IP range, or verification method. Treat the registry token as an unverified deployment signal.

Role and policy boundary

The current OpenIndex page describes an end-to-end encrypted messaging and discovery service for AI agents. It presents agent registration and discovery, search for other agents, OpenClaw integration, and a command-line workflow for profiles and encrypted messages. This is evidence about the current product surface, not proof that a crawler named OICrawler is actively operated or that every request carrying the registry token belongs to OpenIndex.

The registry records this User-Agent:

configuration / code
OICrawler/Nutch https://openindex.ai

The token combines an OICrawler label with Nutch, but the reviewed OpenIndex page does not publish it as a current crawler header, explain its deployment, or identify a controlled crawler network. A matching request could be a legacy deployment, an OpenIndex integration, a private Nutch installation, a test client, or a spoofed value. Do not merge it with all Nutch traffic or attribute it to OpenIndex without source and network evidence.

The current robots file at https://www.openindex.ai/robots.txt contains:

configuration / code
User-agent: *
Disallow: /login
Disallow: /logout

It has no OICrawler-specific group and no Crawl-delay. The global rules should not be interpreted as an OICrawler permission, a crawler authentication method, or a policy for all OpenIndex-related traffic. Robots.txt is advisory and cannot secure user accounts, private messages, keys, or other sensitive data; use authentication, authorization, signed requests, and origin controls.

If logs establish a current, attributable token and you decide to exclude it, use the exact observed token in a deliberate site-owner rule:

configuration / code
User-agent: CONFIRMED-OICRAWLER
Disallow: /

For selective public access after confirmation:

configuration / code
User-agent: CONFIRMED-OICRAWLER
Allow: /public/
Allow: /docs/
Disallow: /login
Disallow: /logout
Disallow: /account/
Disallow: /private/
Disallow: /api/

These are explanatory site-owner controls, not recovered OpenIndex instructions. Protect private agent profiles, encrypted messages, credentials, and APIs at the application layer.

Layered verification

Start with raw access logs and preserve the complete User-Agent, source IP, ASN, reverse-DNS result, HTTP method, requested path, response status, response size, redirect chain, timestamp, concurrency, and request rate. No current OpenIndex source in this review published OICrawler IP ranges, reverse-DNS verification, crawl rate, scheduling, removal contact, or a crawler-specific policy. The registry token is useful for triage but is not authentication.

The OpenIndex hostname in the header is not proof that the request came from OpenIndex. Check source IP and DNS evidence independently, require an authoritative verification path before creating a trust exception, and distinguish a current service endpoint from a crawler identity. A request to public agent-discovery pages may be related to the service, but private messages, login flows, account data, API keys, and encrypted material require authorization regardless of User-Agent.

Compare observed traffic with an agent-discovery hypothesis without turning it into attribution. Public HTML, profile metadata, links, and static assets may be requested by many clients. Private routes, authenticated endpoints, message retrieval, key material, high concurrency, repeated retries, or unexpected bulk downloads establish impact and load risk, not OpenIndex ownership or downstream AI use.

Evaluate /robots.txt independently. Confirm that the intended host serves a successful text response and that the exact OICrawler group you intend to publish is present. The reviewed current file is on www.openindex.ai after redirect and has only global login/logout exclusions. Do not claim that its rules govern a third-party site or authenticate the registry token.

Page-level directives can express indexing preferences:

configuration / code
<meta name="robots" content="noindex, nofollow">
configuration / code
X-Robots-Tag: noindex, nofollow

These signals do not authenticate OICrawler or secure encrypted messages and account resources. Enforce sensitive boundaries in the application and at the origin. If your policy distinguishes agent discovery, search indexing, AI input, reference use, and model training, document each purpose separately; the current OpenIndex product description does not decide downstream use for an individual request.

WAF and Nginx remediation examples

When the identity is unverified, use report-only logging and a narrow observation. Avoid broad OpenIndex, Nutch, agent, or browser matches that can affect unrelated clients:

configuration / code
{
  "description": "Observe unverified OICrawler candidates",
  "expression": "lower(http.user_agent) contains \"oicrawler/nutch\"",
  "action": "log"
}

After independent verification and a policy decision, scope enforcement to sensitive routes and preserve evidence for the rule:

configuration / code
map $http_user_agent $block_confirmed_oicrawler_private {
    default 0;
    # Add only a complete, independently verified token here.
    # ~*OICrawler/Nutch\ https://openindex\.ai 1;
}

server {
    location ~ ^/(login|logout|account|private|messages|keys|api)/ {
        if ($block_confirmed_oicrawler_private) { return 403; }
        try_files $uri $uri/ =404;
    }
}

A User-Agent match is easy to spoof and the token contains a public URL. Do not invent an OpenIndex IP allowlist, ASN exception, reverse-DNS suffix, rate, training policy, or permanent trust rule. Use authentication, authorization, per-client rate limits, concurrency ceilings, timeouts, response-size controls, caching, and anomaly detection at the edge or origin. Start in report-only mode and restrict only after evidence supports the action.

Test public agent pages, static assets, feeds, sitemaps, login/logout, messages, key-bearing endpoints, APIs, 429 behavior, response-size limits, and approved integrations separately. Pair WAF controls with application authorization; robots.txt is not an access-control mechanism.

Review checklist

Search logs for the exact OICrawler/Nutch substring and preserve the complete header, source IP, ASN, PTR result, forward lookup, path, method, response size, status, timing, rate, concurrency, and redirects. Treat openindex.ai in the header as a clue, not proof of ownership.

Re-check the current OpenIndex site, its redirected chat host, robots.txt, official source links, and any future crawler documentation. The reviewed product page describes AI-agent discovery and encrypted messaging but does not publish an OICrawler specification. Keep the profile at documented-limit until a current source establishes the token, purpose, robots behavior, source-verification method, rate guidance, or opt-out path.

Check your robots file for an exact OICrawler group and test precedence and path behavior. Do not generalize OpenIndex’s global /login and /logout exclusions to third-party sites or to all Nutch deployments. Use authentication and authorization for private agent data.

Review agent discovery, search indexing, AI input, reference use, and model-training decisions separately. Neither the registry token nor the current product page establishes downstream permission or use for a specific request. Do not claim successful blocking or verification from configuration alone; validate later logs and obtain an operator response when possible.

Decide whether your objective is to preserve public discovery, reduce load, limit extraction, protect private agent data, or prevent unauthorized message/API access. Publish exact robots rules only after the token is known and enforce sensitive routes at the application and origin layers.

References

  1. OpenIndex — current product entry point; headless-browser review redirected to https://chat.openindex.ai/ and documented end-to-end encrypted messaging and discovery for AI agents on 2026-08-25.
  2. OpenIndex robots.txt — redirected to https://www.openindex.ai/robots.txt and returned a global group disallowing /login and /logout, with no OICrawler group or Crawl-delay, reviewed on 2026-08-25.
  3. OpenIndex CLI repository — linked from the current product page; product source context does not authenticate the registry token or a crawler network.
  4. Crawler User Agents community registry — registry context for the OICrawler/Nutch label; it does not establish current operator identity.
  5. Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
  6. RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.