Bot directory / search-engine

woriobot: Robots.txt & Crawl Policy Reference

Technical reference for the woriobot registry label, noting the absence of current operator documentation, an exact IP list, and verification methods.

AI Summary: The registry describes woriobot as a Worio search engine web crawler but provides no active official documentation URL. Headless-browser research yielded no current operator, crawler specification, IP range, or robots policy for Worio or a "woriobot" crawler. Treat the token as an unverified historical registry label rather than an authenticated active crawler.

Role and policy boundary

The registry describes this entry as a Worio search engine web crawler bot and records the historical User-Agent as Mozilla/5.0 (compatible; woriobot +http://worio.com). The registry provides only a link to the community crawler-user-agents repository rather than an official operator website, and the Worio.com domain is no longer active as a search engine.

During headless-browser review, extensive searches for an official Worio search engine crawler or webmaster guidelines yielded no results. No current operator publishes a crawler policy, explains its purpose, or documents an exact IP list for this label.

This profile therefore uses legacy-label. A matching request could be a legacy deployment, a third-party scraper, a test client, or a spoofed header. Do not infer that an independent Worio search engine operates under this token, that it obeys robots.txt, or that the token establishes permission for AI input, model training, or data retention.

Because no exact current policy was verified, do not publish an assumed woriobot robots group as if it came from the operator. If logs later establish a current, attributable token and you decide to exclude it, use the exact observed value in a deliberate site-owner rule:

configuration / code
User-agent: CONFIRMED-OBSERVED-WORIOBOT
Disallow: /

For selective access after confirmation:

configuration / code
User-agent: CONFIRMED-OBSERVED-WORIOBOT
Allow: /public/
Allow: /docs/
Disallow: /admin/
Disallow: /account/
Disallow: /private/
Disallow: /licensed/
Disallow: /api/

These are explanatory site-owner controls, not recovered operator instructions. Robots.txt is advisory and cannot secure private or licensed content; use authentication, authorization, signed URLs, data minimization, and origin controls.

Layered verification

Start with raw access logs and preserve the complete User-Agent, source IP, ASN, reverse-DNS result, HTTP method, requested path, response status, response size, redirect chain, timestamp, concurrency, and request rate. No current source in this review published a woriobot IP range, reverse-DNS procedure, Crawl-delay, rate guidance, removal address, or verification workflow. The registry label is useful for triage but is not authentication.

Do not treat the word woriobot in a header or a generic registry description as proof of ownership. Check source IP and DNS evidence independently, and require a current authoritative response before creating a trust exception.

Compare observed behavior with a search-indexing hypothesis without turning it into attribution. Public HTML, metadata, feeds, sitemaps, and ordinary assets may be requested by many tools. Private endpoints, licensed content, account routes, APIs, high concurrency, repeated retries, or unexpected bulk downloads establish impact and load risk, not woriobot identity or downstream use.

Evaluate /robots.txt only after the actual observed token is known. Confirm that it is served by the intended host, returns a successful text response, and contains the exact group you intend to publish. An absent operator site is not evidence of robots compliance. If no exact group exists, a global rule may affect unrelated clients and should be adopted only as an explicit site-wide decision.

Page-level directives can express indexing preferences:

configuration / code
<meta name="robots" content="noindex, nofollow">
configuration / code
X-Robots-Tag: noindex, nofollow

These signals do not authenticate an undocumented crawler or secure private paths. Enforce sensitive boundaries in the application and at the origin. If your policy distinguishes search indexing, AI input, reference use, and model training, document each purpose separately rather than inferring permission from a registry label.

WAF and Nginx remediation examples

When the identity is unverified, use report-only logging and a narrow observation. Avoid broad worio or browser rules that can affect unrelated clients:

configuration / code
{
  "description": "Observe unverified woriobot candidates",
  "expression": "lower(http.user_agent) contains \"woriobot\"",
  "action": "log"
}

After independent verification and a policy decision, scope enforcement to sensitive routes and preserve evidence for the rule:

configuration / code
map $http_user_agent $block_confirmed_woriobot_private {
    default 0;
    # Add only a complete, independently verified token here.
    # ~*Exact-Observed-Token 1;
}

server {
    location ~ ^/(admin|account|private|licensed|internal|api)/ {
        if ($block_confirmed_woriobot_private) { return 403; }
        try_files $uri $uri/ =404;
    }
}

A User-Agent match is easy to spoof, and the registry token is undocumented. Do not invent an IP allowlist, ASN exception, reverse-DNS suffix, rate, training policy, or permanent trust exception. Use per-client rate limits, concurrency ceilings, timeouts, response-size controls, caching, and anomaly detection at the edge or origin. Start in report-only mode, review false positives, and restrict only after evidence supports the action.

Test public pages, feeds, sitemaps, structured data, licensed assets, account routes, APIs, 429 behavior, response-size limits, and approved integrations separately. Pair WAF controls with authentication and application authorization instead of using robots.txt as an access-control mechanism.

Review checklist

Search logs for the exact woriobot substring and preserve the complete header, source IP, ASN, PTR result, forward lookup, path, method, response size, status, timing, rate, concurrency, and redirects. The registry does not provide a verification method.

Re-check the crawler registry, any future operator source, robots.txt, and operator contact path. In this review no official operator documentation could be located. Keep this profile at legacy-label until a current source publishes an exact token, purpose, robots behavior, source-verification method, rate guidance, or opt-out process.

Decide whether your objective is to preserve public discovery, reduce crawl load, limit extraction, protect licensed material, or prevent private access. Publish exact robots rules only after the token is known; enforce private routes with authentication and origin controls.

Review search indexing, AI input, reference use, and model-training decisions separately. The registry description alone does not establish downstream permission or use. Do not claim successful blocking or verification from configuration alone; validate later logs and obtain a current operator response when possible.

Record the date and reason for the legacy-label classification so future evidence can be compared without silently upgrading an unsupported historical label. If a new deployment identifies itself differently, create a separate evidence trail.

References

  1. Crawler User Agents community registry — registry context for the woriobot label; it does not authenticate current infrastructure or provide an active documentation URL.
  2. Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
  3. RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.