Bot directory / search-engine

jyxobot: Robots.txt & Crawl Policy Reference

Technical reference for the historical Jyxo search-bot registry label, with explicit limits around its undocumented User-Agent and current Cloudflare host timeout.

AI Summary: jyxobot is a historical registry label associated with Jyxo search crawling, but no exact User-Agent is published in the available inventory. Both https://jyxo.cz/ and https://jyxo.cz/robots.txt returned Cloudflare 522 Connection timed out pages during review, so current policy, activity, source network, and verification could not be established.

Role and policy boundary

The registry describes jyxobot as a Jyxo search-engine bot for web crawling, but it supplies no User-Agent value. A likely operator domain, https://jyxo.cz/, was tested with the headless browser and returned a Cloudflare 522 response stating that the browser and Cloudflare were working while the jyxo.cz host did not respond. The direct robots path returned the same 522 page.

These observations do not establish a current bot identity, current operator, or policy. The registry’s search-crawling description is a historical role hypothesis only. A request may come from an old deployment, a fork, an independent tool, or a spoofed header. Do not infer current activity, operator identity, downstream use, AI-training purpose, permission to crawl, or access to private content from the label or the domain timeout.

Because the exact token is unknown, do not publish a made-up User-agent: jyxobot group as if it will reliably match the client. If logs later reveal a complete and independently attributable header, use that exact value:

configuration / code
User-agent: CONFIRMED-OBSERVED-TOKEN
Disallow: /

For selective access after confirmation:

configuration / code
User-agent: CONFIRMED-OBSERVED-TOKEN
Allow: /public/
Allow: /docs/
Disallow: /admin/
Disallow: /account/
Disallow: /private/
Disallow: /api/

These are defensive examples, not recovered Jyxo instructions. Robots.txt is advisory and cannot protect private or licensed content; use authentication, authorization, signed URLs, and origin controls for those boundaries.

Layered verification

Begin with raw access logs and preserve the complete User-Agent, source IP, ASN, reverse-DNS result, HTTP method, requested path, response status, response size, redirect chain, timestamp, and request rate. No exact jyxobot token or current Jyxo source-verification procedure was available in this review. Do not reuse rules for Googlebot, another search crawler, or an unrelated Czech service.

Classify behavior by observable impact, not by a guessed name. Public HTML, canonical metadata, feeds, and sitemaps may be consistent with indexing. Private endpoint access, high concurrency, repeated retries, unexpected downloads, or traffic that ignores your published restrictions may indicate spoofing, abuse, a fork, or another client. These observations establish operational risk, not operator identity or downstream use.

Evaluate /robots.txt only after the actual header is known. Confirm that it is served by the correct host, returns a successful status and text content type, and contains the exact group you intend to publish. The current 522 response is not a valid robots policy and must not be interpreted as allow or disallow evidence. Page-level directives can express discovery preferences:

configuration / code
<meta name="robots" content="noindex, nofollow">
configuration / code
X-Robots-Tag: noindex, nofollow

These signals do not authenticate an undocumented Jyxo client or secure private routes. Enforce sensitive boundaries in the application and at the origin.

WAF and Nginx remediation examples

When the identity is unknown, use a report-only candidate rule and do not block broad Jyxo or bot substrings. Adapt the expression to your WAF provider:

configuration / code
{
  "description": "Log unidentified Jyxo crawler candidates",
  "expression": "lower(http.user_agent) contains \"jyxo\"",
  "action": "log"
}

Once logs and an authoritative source establish a complete token, replace the candidate with a narrow route-scoped control:

configuration / code
map $http_user_agent $block_confirmed_jyxo {
    default 0;
    # Add only a complete, independently verified token here.
    # ~*Confirmed-Jyxo-Token 1;
}

server {
    location ~ ^/(admin|account|private|internal|api)/ {
        if ($block_confirmed_jyxo) { return 403; }
        try_files $uri $uri/ =404;
    }
}

Do not invent an IP allowlist, reverse-DNS suffix, or permanent deny rule from a Cloudflare 522. A User-Agent is easy to spoof and the candidate expression may catch unrelated clients. Test public HTML, feeds, sitemaps, media, accounts, APIs, and approved integrations separately. Pair edge controls with authentication, rate limits, signed assets, caching, and anomaly detection.

Review checklist

Search logs for jyxo only as a candidate, then preserve the complete request header. Record source IP, ASN, PTR result, forward lookup, path, method, response size, status, timing, rate, and redirect chain. Check whether a current operator source identifies the client before assigning a purpose or robots rule.

Re-check jyxo.cz, its robots path, and the community registry when the host becomes reachable. Keep this profile at legacy-label and the User-Agent as Not publicly documented until a credible source publishes a current token, purpose, source-verification method, or robots guidance. Treat the 522 result as a host-availability limitation, not proof of inactivity.

Decide whether your objective is to preserve search visibility, limit extraction, protect private material, or reduce crawl load. Publish exact robots rules only after the token is known, and enforce private routes with application and WAF controls. Do not claim successful blocking from configuration alone; verify subsequent logs and response behavior.

References

  1. Jyxo homepage — likely operator domain; headless-browser request returned Cloudflare 522 Connection timed out on 2026-08-25.
  2. Jyxo robots.txt — headless-browser request returned the same Cloudflare 522 Connection timed out; no directives were read on 2026-08-25.
  3. Crawler User Agents community registry — registry source for the label; no exact jyxobot User-Agent was available in the inventory.
  4. Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
  5. RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.