Bot directory / search-engine

purebot: Robots.txt & Crawl Policy Reference

Technical reference for the historical purebot registry label, with explicit limits because no exact User-Agent or current operator documentation was verified.

AI Summary: The registry labels purebot as a web crawler for content discovery but provides no exact User-Agent and no operator documentation. A headless-browser search for an official purebot source and robots.txt returned no relevant first-party evidence. Current purpose, identity, robots behavior, rate, IP range, reverse-DNS verification, and opt-out are unverified. Treat the label as historical context rather than an authenticated crawler identity.

Role and policy boundary

The registry describes purebot as a pure web crawler for content discovery, but it does not provide an exact User-Agent. This profile deliberately records Not publicly documented rather than inventing a token from the name. A headless-browser query for purebot web crawler official robots.txt returned no readable relevant first-party source, so the registry description cannot be upgraded to a current operator policy.

No current operator, deployment, product, search index, content-analysis purpose, or service lifecycle was established. A request that an analyst informally calls “purebot” may be an unrelated client, a private tool, a legacy service, a test harness, or a spoofed label. The name does not establish search indexing, AI input, model training, data retention, permission to crawl, or access to private content.

Because no exact current token was verified, do not publish a purebot robots group as if it came from an operator. First identify the actual header in logs. If it is a confirmed token and your site policy is to exclude it, use that exact observed value:

configuration / code
User-agent: CONFIRMED-OBSERVED-TOKEN
Disallow: /

For selective access after confirmation:

configuration / code
User-agent: CONFIRMED-OBSERVED-TOKEN
Allow: /public/
Allow: /docs/
Disallow: /admin/
Disallow: /account/
Disallow: /private/
Disallow: /licensed/
Disallow: /api/

These are explanatory site-owner controls, not recovered purebot instructions. Robots.txt is advisory and cannot secure private or licensed content; use authentication, authorization, signed URLs, data minimization, and origin controls.

Layered verification

Start with raw access logs and preserve the complete User-Agent, source IP, ASN, reverse-DNS result, HTTP method, requested path, response status, response size, redirect chain, timestamp, concurrency, and request rate. No exact purebot token, current operator source, IP range, reverse-DNS procedure, Crawl-delay, rate guidance, removal address, or verification workflow was established in this review.

Do not classify a request from a path, hostname, reverse-DNS guess, or informal bot name alone. Check the complete header and source network independently. A User-Agent match is not authentication, and without an exact token there is no safe bot-specific allowlist or block expression to copy from this profile.

Compare observed behavior with the historical content-discovery hypothesis without turning it into attribution. Public HTML, metadata, feeds, sitemaps, and ordinary assets may be requested by many tools. Private endpoints, licensed material, account routes, APIs, high concurrency, repeated retries, or unexpected bulk downloads establish impact and load risk, not purebot identity or downstream use.

Evaluate /robots.txt only after the actual observed token is known. Confirm that it is served by the intended host, returns a successful text response, and contains the exact group you intend to publish. No policy can be attributed to purebot from the registry label or a missing source. If no exact group exists, a global rule may affect unrelated clients and should be adopted only as an explicit site-wide decision.

Page-level directives can express indexing preferences:

configuration / code
<meta name="robots" content="noindex, nofollow">
configuration / code
X-Robots-Tag: noindex, nofollow

These signals do not authenticate an undocumented crawler or secure private paths. Enforce sensitive boundaries in the application and at the origin. If your policy distinguishes search indexing, AI input, reference use, and model training, document each purpose separately rather than inferring permission from a vague registry label.

WAF and Nginx remediation examples

When no exact token is known, do not write a production rule matching purebot or every generic crawler. Use report-only logging of the complete request and investigate manually:

configuration / code
{
  "description": "Review suspected purebot traffic without attribution",
  "expression": "true",
  "action": "log",
  "fields": ["http.user_agent", "source.ip", "request.path", "response.status", "request.rate"]
}

After an exact token and authoritative identity are established, replace the placeholder with a narrow, route-scoped rule:

configuration / code
map $http_user_agent $block_confirmed_purebot_private {
    default 0;
    # Add only a complete, independently verified token here.
    # ~*Exact-Observed-Token 1;
}

server {
    location ~ ^/(admin|account|private|licensed|internal|api)/ {
        if ($block_confirmed_purebot_private) { return 403; }
        try_files $uri $uri/ =404;
    }
}

The example intentionally does not invent an identity. Do not infer an IP allowlist, reverse-DNS suffix, rate, training policy, or permanent trust exception. Use per-client rate limits, concurrency ceilings, timeouts, response-size controls, caching, and anomaly detection at the edge or origin. Start in report-only mode, review false positives, and restrict only after evidence supports the action.

Test public pages, feeds, sitemaps, structured data, licensed assets, account routes, APIs, 429 behavior, response-size limits, and approved integrations separately. Pair WAF controls with authentication and application authorization rather than using robots.txt as an access-control mechanism.

Review checklist

Search logs for suspected crawler traffic without assuming that a string purebot is present. Preserve the complete User-Agent, source IP, ASN, PTR result, forward lookup, path, method, response size, status, timing, rate, concurrency, and redirects. The registry does not provide an exact token.

Re-check the crawler registry, any future purebot operator source, robots.txt, and contact path. The headless search reviewed here did not establish a relevant first-party source. Keep this profile at legacy-label until a current source publishes an exact token, purpose, robots behavior, source-verification method, rate guidance, or opt-out process.

Decide whether your objective is to preserve public discovery, reduce crawl load, limit extraction, protect licensed material, or prevent private access. Publish exact robots rules only after the token is known; enforce private routes with authentication and origin controls.

Review search indexing, AI input, reference use, and model-training decisions separately. Neither the registry label nor an unknown User-Agent establishes downstream permission or use. Do not claim successful blocking or verification from configuration alone; validate later logs and obtain an operator response when possible.

Record the date and reason for the legacy-label classification. If a future deployment identifies itself with a concrete token, create a new evidence trail and update policy only after confirming that source and network.

References

  1. Crawler User Agents community registry — registry context for the purebot label; no exact User-Agent or current operator was established in this review.
  2. Headless Bing search for purebot — reviewed on 2026-08-25; no relevant first-party source was established and snippets were not treated as evidence.
  3. Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
  4. RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.