Bot directory / ai-search

ZanistaBot: Robots.txt & Crawl Policy Reference

Technical reference for the ZanistaBot registry label associated with Zanista.AI. Learn how to verify its identity when the first-party crawler page does not expose a usable policy.

AI Summary: The registry associates ZanistaBot/1.0 with Zanista.AI and an official crawler-info URL, but the first-party page reviewed rendered only a loading shell and exposed no usable crawler contract. The token, source ranges, purpose, crawl rate, and robots behavior are therefore unverified. Treat observed traffic as evidence to investigate and use layered access controls.

Role and policy boundary

The registry describes ZanistaBot as Zanista's AI search crawler. The first-party domain is active and exposes Zanista.AI product navigation, including Products and PaperPal, but the crawler-info route did not reveal a policy document during this review. A product association and a URL in a User-Agent are not proof that a request is currently operated by Zanista.AI.

The registry header is:

configuration / code
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; ZanistaBot/1.0; +https://zanista.ai/crawler-info) Chrome/W.X.Y.Z Safari/537.36

No operator-published source IP ranges, crawl schedule, rate limit, robots statement, or verification method was confirmed. Do not infer that ZanistaBot is collecting training data, building a search index, or respecting robots.txt solely from this label. Keep the versioned token and the stable ZanistaBot name separate when classifying logs.

If your logs confirm the exact token and you want to communicate a restriction, publish:

configuration / code
User-agent: ZanistaBot
Disallow: /

For selective access:

configuration / code
User-agent: ZanistaBot
Allow: /docs/
Allow: /public-reference/
Disallow: /private/
Disallow: /internal/
Disallow: /api/

Robots.txt is advisory and cannot protect private, paywalled, or licensed content. Use authentication and authorization for those boundaries.

Layered verification

Start with raw access logs and preserve the full User-Agent, source IP, ASN, reverse DNS, method, path, status, response size, redirects, timestamp, and request rate. The header is self-declared and browser-like; the linked URL is not a source-authentication mechanism.

Analyze behavior without assigning purpose prematurely. Requests for public documents, feeds, sitemaps, and search-oriented metadata may be consistent with discovery, while high-concurrency traversal, repeated retries, original-media downloads, or private API access may indicate scraping or abuse. These signals establish operational impact but cannot prove vendor identity or downstream AI use.

Evaluate /robots.txt independently. Confirm the canonical host, status, content type, exact group, and path match. During the review, the first-party crawler page exposed no usable robots policy, so an allowed request is not evidence of Zanista compliance. Page-level metadata may express a discoverability preference:

configuration / code
<meta name="robots" content="noindex, nofollow">
configuration / code
X-Robots-Tag: noindex, nofollow

These directives do not secure private routes and may not be honored by an undocumented client. Use authenticated delivery, signed URLs, and origin controls.

WAF and Nginx remediation examples

If logs confirm unwanted requests carrying the token, use a narrow WAF rule and monitor false positives:

configuration / code
{
  "description": "Block observed ZanistaBot token",
  "expression": "lower(http.user_agent) contains \"zanistabot\"",
  "action": "block"
}

For Nginx, scope enforcement to private and high-cost routes while investigating public access:

configuration / code
map $http_user_agent $block_zanistabot {
    default 0;
    ~*ZanistaBot 1;
}

server {
    location ~ ^/(private|internal|account|paywall|uploads|api)/ {
        if ($block_zanistabot) { return 403; }
        try_files $uri $uri/ =404;
    }
}

A User-Agent match is easy to spoof or evade and may block an approved integration. Do not create an IP allowlist without a verified operator-published range. Test browsers, social previews, feed readers, search crawlers, and monitors. Pair edge matching with authentication, rate limits, signed assets, and anomaly detection.

Review checklist

Search logs for ZanistaBot and preserve the complete header, source IP, paths, response sizes, statuses, timing, and concurrency. Re-check the first-party crawler-info URL after changes; during this review it showed a loading shell without crawler policy text. Keep the registry's AI-search description labeled as unverified.

Decide whether your objective is to preserve potential search visibility, prevent extraction, protect private content, or reduce crawl load. Publish a targeted robots group for the exact observed token, enforce sensitive routes with WAF and application authorization, and test docs, media, feeds, sitemaps, uploads, and APIs separately. Revisit the profile if Zanista publishes a stable User-Agent, source verification, rate policy, or robots statement.

References

  1. Zanista crawler information — first-party URL reviewed on 2026-08-24; it rendered a Zanista.AI shell but no usable crawler policy was exposed.
  2. Zanista.AI — first-party domain and product context; product navigation alone does not verify a crawler contract.
  3. Google Robots.txt Introduction — general explanation of crawler directives and their limitations.

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.