Bot directory / ai-search

Timpibot: Robots.txt & Crawl Policy Reference

Comprehensive guide to Timpibot, the crawler for the decentralized Timpi search engine. Learn how to manage its access and address its aggressive crawling behavior.

AI Summary: Timpibot is the web crawler for Timpi, a decentralized, Web3-integrated search engine project that aims to build an independent web index infrastructure for AI and search. Webmasters have frequently reported aggressive crawling behavior and poor robots.txt compliance from this bot.

Role and policy boundary

Timpibot indexes the web to build Timpi's proprietary data index. This index is intended to power their decentralized search engine and provide data APIs for AI workflows (such as their ORCA AI orchestration).

As a newer, alternative crawler, webmasters must evaluate whether the potential traffic from a decentralized search engine justifies the server resources consumed by the bot. Given reports of it ignoring directives and sending dozens of parallel requests per second, many administrators opt for a default-deny stance.

A robots rule is a declaration of intent; it does not replace authentication, authorization, or rate limiting. Start with a dedicated group:

configuration / code
User-agent: Timpibot
Allow: /
Disallow: /staging/
Disallow: /internal/

To request the bot to stop accessing the entire site, use:

configuration / code
User-agent: Timpibot
Disallow: /

Avoid assuming that User-agent: * expresses the same business intent. A wildcard can affect assistant and training crawlers too, and it makes later audits harder because the source of the decision is less specific.

Layered verification

Verify the same URL through each control plane instead of assuming that one green signal represents the whole request path. Compare the bot-specific robots group, the page-level metadata, and the response headers captured at the public edge.

There are widespread industry reports regarding Timpibot's lack of adherence to robots.txt. While they should theoretically respect the Timpibot User-Agent directive, relying solely on robots.txt is often insufficient for this specific crawler.

The Policy Engine evaluates the selected user-agent, path scope, and the other supplied layers independently. It can therefore explain why a bot is allowed while another is blocked, rather than returning one blended website score.

Page-level directives can still override the intended outcome for indexing:

configuration / code
<meta name="robots" content="noindex, nofollow">
configuration / code
X-Robots-Tag: noindex, nofollow

If a response uses these tags, the report marks the result as blocked or conflicting even if the crawler-specific robots group is permissive. This is especially important for canonical pages served through an edge cache where headers may differ from the origin response.

WAF and Nginx remediation examples

Due to reports of highly aggressive crawling and robots.txt non-compliance, blocking Timpibot at the WAF or server level is strongly recommended if you do not want your site indexed by their service or if you are experiencing performance degradation.

configuration / code
{
  "description": "Block Timpibot",
  "expression": "lower(http.user_agent) contains \"timpibot\"",
  "action": "block"
}

Use your platform's actual middleware response pattern rather than copying this simplified example without review. Never place a secret, verification token, or internal policy identifier in a public header.

configuration / code
map $http_user_agent $block_timpibot_private {
    default 0;
    ~*Timpibot 1;
}

server {
    location ~ ^/(admin|account|private|licensed|internal|api)/ {
        if ($block_timpibot_private) { return 403; }
        try_files $uri $uri/ =404;
    }
}

Review checklist

Use this checklist after every policy change and after a CDN or WAF migration. Record the request URL, User-Agent, HTTP status, final redirect, and the exact evidence used to reach the decision.

Verify that the dedicated group appears before relying on a wildcard, test a representative public and private path, and compare live response headers with robots.txt. Keep the policy close to the content owner's intent and record whether the site wants discovery, citation, or no access at all.


Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.