Bot directory / search-engine

trovitBot: Robots.txt & Crawl Policy Reference

Technical reference for trovitBot, the web crawler operated by Trovit, including its robots.txt compliance, crawl rate, and content discovery purpose.

AI Summary: trovitBot is the official web crawler operated by Trovit to discover and index new and updated pages. The official documentation confirms it respects the trovitBot robots.txt token and aims to access sites no more than once per second. Because the operator does not publish a static IP list or a specific reverse-DNS verification method, use standard network validation and behavioral analysis before trusting the header.

Role and policy boundary

Trovit operates trovitBot to crawl the web, discover new and updated pages, and add them to its search index. The registry records this User-Agent string:

configuration / code
Mozilla/5.0 (compatible; trovitBot 1.0; +http://www.trovit.com/bot.html)

The operator explicitly states that TrovitBot respects the standard robots.txt protocol and identifies itself using the trovitBot token. For selective access to public pages:

configuration / code
User-agent: trovitBot
Allow: /public/
Allow: /jobs/
Disallow: /admin/
Disallow: /account/
Disallow: /private/
Disallow: /api/

For a complete exclusion from the Trovit index:

configuration / code
User-agent: trovitBot
Disallow: /

These are site-owner controls. Robots.txt is advisory and cannot secure private or licensed content; use authentication, authorization, signed URLs, data minimization, and origin controls for those boundaries. The search-indexing purpose does not establish permission for generative AI answers or third-party model training.

Layered verification

Preserve the complete User-Agent, source IP, ASN, reverse-DNS result, HTTP method, requested path, response status, response size, redirect chain, timestamp, concurrency, and request rate in logs. The reviewed Trovit documentation does not publish a static IP list or a specific reverse-DNS verification procedure.

Because a User-Agent is easily spoofed, do not treat the header alone as authentication. Check source IP and DNS evidence independently. A trovitBot header from a residential network, a known proxy, or a cloud provider without an established Trovit affiliation should not be treated as authentic traffic.

Compare observed traffic with the documented search-indexing role. Public HTML, metadata, feeds, sitemaps, and ordinary assets are consistent with discovery. Private endpoints, account routes, APIs, licensed content, high concurrency, repeated retries, or unexpected bulk downloads establish impact and load risk, not permission. The operator states that, on average, TrovitBot accesses most sites no more than once per second (though network lag may occasionally cause brief spikes).

Evaluate /robots.txt independently. Confirm that your host returns a successful text response and contains the exact trovitBot group you intend to publish. Check precedence, path matching, and actual request behavior. Meta directives can express indexing preferences:

configuration / code
<meta name="robots" content="noindex, nofollow">
configuration / code
X-Robots-Tag: noindex, nofollow

These signals do not authenticate the crawler or protect private routes. Enforce sensitive boundaries in the application and at the origin.

WAF and Nginx remediation examples

Begin in report-only mode and correlate the trovitBot User-Agent with independent network verification. Adapt the expression to your WAF provider:

configuration / code
{
  "description": "Observe trovitBot candidates before enforcement",
  "expression": "lower(http.user_agent) contains \"trovitbot\"",
  "action": "log"
}

For a deliberately restricted private route, use the exact token only after verifying the source network:

configuration / code
map $http_user_agent $block_trovit_private {
    default 0;
    ~*trovitBot 1;
}

server {
    location ~ ^/(admin|account|private|licensed|internal|api)/ {
        if ($block_trovit_private) { return 403; }
        try_files $uri $uri/ =404;
    }
}

This rule is route-scoped and does not prove identity. Do not invent an IP allowlist, Trovit ASN exception, rate, or training permission without operator documentation.

A User-Agent is easy to spoof. Use per-client rate limits, concurrency ceilings, timeouts, response-size controls, caching, and anomaly detection at the edge or origin. Start in report-only mode, review false positives, and restrict only after evidence supports the action.

Test public pages, feeds, sitemaps, structured data, licensed assets, account routes, APIs, 429 behavior, response-size limits, and approved integrations separately. Pair WAF controls with authentication and application authorization instead of using robots.txt as an access-control mechanism.

Review checklist

Search logs for the exact trovitBot string, preserving the full header, source IP, ASN, PTR result, forward lookup, path, method, response size, status, timing, rate, concurrency, and redirects.

Verify the source network independently before trusting the client. Treat requests from unverified networks as spoofed.

Check your robots file for an exact trovitBot group. Test precedence and path behavior, and verify actual logs after publishing. Use META or X-Robots-Tag for indexing preferences and authentication for private resources.

Review search indexing, AI input, reference use, and model-training decisions separately. The operator explicitly states the crawler is used for its search index.

Keep the profile at documented-limit: the operator publishes the crawler identity, robots compliance, rate limits, and purpose, but lacks a static IP list or reverse-DNS verification method. Re-check the crawler documentation when the source network or traffic pattern changes. Do not claim successful blocking or verification from configuration alone; validate later logs.

References

  1. TrovitBot Information — official page documenting the trovitBot robots.txt token, crawl rate, and indexing purpose; reviewed with the headless browser on 2026-08-25.
  2. Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
  3. RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.