Bot directory / search-engine

NapBot: Robots.txt & Crawl Policy Reference

Technical reference for the historical NapBot registry label, with explicit limits around the parked napbot.com domain and its current permissive robots file.

AI Summary: NapBot is a historical registry label associated with web-content indexing, but its domain currently redirects to a HugeDomains sale page and no exact User-Agent or operator-owned crawler policy is available. The current napbot.com/robots.txt contains only an empty global Disallow: value; because the domain is parked, this must not be interpreted as a live NapBot permission or policy.

Role and policy boundary

The registry describes NapBot as a web-content indexing and search-engine crawler, but it provides no User-Agent and links only to http://napbot.com/. A headless-browser request to the domain redirected to a HugeDomains domain-profile page stating that NapBot.com is for sale for $15,195. This shows the domain’s current parked/sale state; it does not establish who operated the historical bot or whether a client remains active.

The direct https://napbot.com/robots.txt currently returns:

configuration / code
User-agent: *
Disallow:

An empty global Disallow: value is permissive for the currently served domain, but it is not evidence of an intentional NapBot crawler policy. A parked domain can use registrar defaults, and future ownership or hosting can change the file. Do not infer that the historical NapBot label is allowed to crawl your site, that the crawler is active, or that the domain owner is the historical operator.

The search-indexing role is therefore a registry hypothesis only. A request may come from an old deployment, an independent crawler, a test client, a fork, or a spoofed header. Do not infer current activity, operator identity, downstream use, AI-training purpose, permission to crawl, or access to private content from the label or parked domain.

Because the exact token is unknown, do not publish a made-up User-agent: NapBot group as if it will reliably match the client. If logs later reveal a complete and independently attributable token, use that exact value:

configuration / code
User-agent: CONFIRMED-OBSERVED-TOKEN
Disallow: /

For selective access after confirmation:

configuration / code
User-agent: CONFIRMED-OBSERVED-TOKEN
Allow: /public/
Allow: /docs/
Disallow: /admin/
Disallow: /account/
Disallow: /private/
Disallow: /api/

These are defensive examples, not recovered NapBot instructions. Robots.txt is advisory and cannot protect private or licensed content; use authentication, authorization, signed URLs, data minimization, and origin controls for those boundaries.

Layered verification

Begin with raw access logs and preserve the complete User-Agent, source IP, ASN, reverse-DNS result, HTTP method, requested path, response status, response size, redirect chain, timestamp, and request rate. No exact NapBot token or current operator source-verification procedure was available in this review. A parked domain and permissive global robots file provide no authentication evidence.

Compare observed behavior with a search-indexing hypothesis without turning it into attribution. Public HTML, canonical metadata, feeds, and sitemaps may be requested by many tools. Private endpoints, account routes, licensed documents, APIs, high concurrency, repeated retries, or unexpected bulk downloads may indicate spoofing, abuse, a partner integration, or another client. These observations establish impact and data risk, not operator identity or downstream use.

Evaluate your own /robots.txt independently. Confirm it is served from the intended host, returns a successful status and text content type, and contains the exact group you intend to publish. Never copy the parked napbot.com file as if it were a crawler policy for your site. Test the complete observed token, a confirmed NapBot group if one becomes available, and the global group separately. Page-level directives can express discovery preferences:

configuration / code
<meta name="robots" content="noindex, nofollow">
configuration / code
X-Robots-Tag: noindex, nofollow

These signals do not authenticate an undocumented crawler or secure private routes. Enforce sensitive boundaries in the application and at the origin.

WAF and Nginx remediation examples

When the identity is unknown, use report-only logging and avoid a broad nap or bot substring rule. Adapt the expression to your WAF provider:

configuration / code
{
  "description": "Log unidentified NapBot candidates",
  "expression": "lower(http.user_agent) contains \"napbot\"",
  "action": "log"
}

Once a complete token and authoritative source establish the identity, replace the candidate with a narrow route-scoped control:

configuration / code
map $http_user_agent $block_confirmed_napbot_private {
    default 0;
    # Add only a complete, independently verified token here.
    # ~*Confirmed-NapBot-Token 1;
}

server {
    location ~ ^/(admin|account|private|licensed|internal|api)/ {
        if ($block_confirmed_napbot_private) { return 403; }
        try_files $uri $uri/ =404;
    }
}

Do not invent an IP allowlist, reverse-DNS suffix, or registrar-based exception. A User-Agent match is easy to spoof and a parked-domain robots file is not a reliable policy source. Test public pages, feeds, sitemaps, licensed documents, accounts, APIs, and approved integrations separately. Pair edge controls with authentication, rate limits, signed assets, caching, data-loss monitoring, and anomaly detection.

Review checklist

Search logs for broad napbot candidates only to locate a complete header, then preserve the exact value. Record source IP, ASN, PTR result, forward lookup, path, method, response size, status, timing, rate, and redirect chain. Check whether a current operator source identifies the client before assigning search purpose or a robots rule.

Re-check napbot.com, the domain registration/hosting state, the community registry, and any newly published bot documentation when a credible source becomes available. Keep this profile at legacy-label and the User-Agent as Not publicly documented until a first-party source publishes a current token, purpose, source-verification method, rate guidance, or opt-out process.

Treat the HugeDomains sale page and empty robots file as observations about the current domain, not proof of bot inactivity or permission to crawl. Do not claim successful blocking from configuration alone; verify subsequent access logs and response behavior.

Decide whether your objective is to preserve public discovery, limit extraction, protect private material, or reduce crawl load. Publish exact robots rules only after the token is known, and enforce private routes with application and WAF controls.

References

  1. NapBot.com domain profile — current redirect target; states NapBot.com is for sale for $15,195, observed with the headless browser on 2026-08-25.
  2. NapBot robots.txt — current file containing User-agent: * and an empty Disallow:, observed with the headless browser on 2026-08-25; it was not treated as a historical crawler policy.
  3. Crawler User Agents community registry — registry context for the label; it provides no exact NapBot User-Agent in the inventory.
  4. Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
  5. RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.