← Bot Directory/INETDEX-BOT
Bot directory / search-engine

INETDEX-BOT: Robots.txt & Crawl Policy Reference

Technical reference for the historical INETDEX-BOT indexing label, with explicit limits around unavailable documentation and current request verification.

AI Summary: INETDEX-BOT/1.5 is a historical registry User-Agent associated with INETDEX indexing. The linked documentation failed with net::ERR_CONNECTION_CLOSED over HTTPS and net::ERR_EMPTY_RESPONSE over HTTP during review, so no current operator policy, robots behavior, IP range, or verification method was established. Treat matching traffic as an unverified historical observation.

Role and policy boundary

The registry describes INETDEX-BOT as an INETDEX web crawler for indexing and records this header:

configuration / code
INETDEX-BOT/1.5 (Mozilla/5.0; https://inetdex.com/bot.html)

The linked documentation URL could not be reached during the headless-browser review. The HTTPS request closed the connection, and the HTTP equivalent returned an empty response. These outcomes prevent confirmation of a current operator, production service, robots policy, rate guidance, source network, or opt-out process. They do not prove that the historical service never existed or that every matching request is inactive.

Treat search indexing as the registry’s historical role hypothesis only. A matching header may come from a legacy deployment, an independent client, a fork, a test harness, or a spoofed request. Do not infer current activity, operator identity, downstream use, AI-training purpose, permission to crawl, or access to private content from the token.

If access logs confirm this exact token and your site policy is to exclude it from discovery, a narrow defensive rule could be:

configuration / code
User-agent: INETDEX-BOT
Disallow: /

For selective access to approved public pages:

configuration / code
User-agent: INETDEX-BOT
Allow: /public/
Allow: /docs/
Disallow: /admin/
Disallow: /account/
Disallow: /private/
Disallow: /api/

These are site-owner examples, not recovered INETDEX instructions. Robots.txt is advisory and cannot protect private or licensed content; use authentication, authorization, signed URLs, and origin controls for those boundaries.

Layered verification

Begin with raw access logs and preserve the complete User-Agent, source IP, ASN, reverse-DNS result, HTTP method, requested path, response status, response size, redirect chain, timestamp, and request rate. No reachable current INETDEX source published a network range, reverse-DNS procedure, rate policy, or canonical contact. The registry header alone is not authentication.

Compare observed behavior with a search-indexing hypothesis without turning it into attribution. Requests for public HTML, canonical metadata, feeds, sitemaps, and ordinary assets may be consistent with indexing. Private endpoint access, high concurrency, repeated retries, unexpected downloads, or traffic that ignores your policy may indicate spoofing, abuse, a fork, or a different client. These observations establish impact, not operator identity or downstream use.

Evaluate /robots.txt independently if you choose to publish a rule. Confirm that it is served from the intended host, returns a successful status and text content type, and contains the exact group you intend to apply. Test the complete token and any global group separately. Page-level directives can express discovery preferences:

configuration / code
<meta name="robots" content="noindex, nofollow">
configuration / code
X-Robots-Tag: noindex, nofollow

These signals do not authenticate a historical INETDEX client or secure private routes. Enforce sensitive boundaries in the application and at the origin.

WAF and Nginx remediation examples

If logs show a repeatable unwanted token, start in report-only mode and preserve representative requests. Adapt the expression to your WAF provider; it identifies a historical header pattern only:

configuration / code
{
  "description": "Review historical INETDEX-BOT traffic",
  "expression": "lower(http.user_agent) contains \"inetdex-bot\"",
  "action": "log"
}

After reviewing false positives and confirming the business decision, scope enforcement to sensitive routes:

configuration / code
map $http_user_agent $review_inetdex_bot {
    default 0;
    ~*INETDEX-BOT 1;
}

server {
    location ~ ^/(admin|account|private|internal|api)/ {
        if ($review_inetdex_bot) { return 403; }
        try_files $uri $uri/ =404;
    }
}

Do not invent an IP allowlist, reverse-DNS suffix, or current operator exception. A User-Agent is easy to spoof and a broad historical match can catch an authorized test or unrelated client. Test public HTML, feeds, sitemaps, media, uploads, account flows, APIs, and approved integrations separately. Pair edge controls with authentication, rate limits, signed assets, caching, and anomaly detection.

Review checklist

Search logs for the complete registered header and preserve representative source IPs, ASNs, reverse-DNS results, paths, methods, response sizes, statuses, timing, and rate. Check whether the source can be independently verified and whether observed behavior resembles public search indexing. Do not classify a request as authentic from the token or registry entry.

Re-check inetdex.com, the /bot.html path, and the registry when the service becomes reachable. Keep this profile at legacy-label until a credible operator publishes a current policy, canonical User-Agent, source-verification method, or robots guidance. Treat the HTTPS and HTTP failures as source-access limitations, not proof of current inactivity.

Decide whether your objective is to preserve search visibility, limit extraction, protect private material, or reduce crawl load. Publish exact robots rules and enforce private routes with application and WAF controls. Do not claim successful blocking from configuration alone; verify subsequent access logs and response behavior.

References

  1. INETDEX-BOT documentation — registry-linked source; direct HTTPS headless-browser request failed with net::ERR_CONNECTION_CLOSED on 2026-08-25.
  2. INETDEX-BOT documentation HTTP equivalent — direct HTTP headless-browser request returned net::ERR_EMPTY_RESPONSE on 2026-08-25.
  3. Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
  4. RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.