Bot directory / ai-training

Thinkbot: Robots.txt & Crawl Policy Reference

Technical reference for the experimental Thinkbot User-Agent observed in webmaster logs. Learn why its operator and purpose remain unverified.

AI Summary: Thinkbot is an experimental User-Agent observed in a webmaster report, not a currently verified operator policy. The report records a long self-descriptive header, no robots.txt checks, and traffic from many addresses, while saying a possible thinkbot.agency association could not be confirmed. Treat the token and network observations as local evidence only; use log-first investigation and active controls when necessary.

Role and policy boundary

The registry categorizes Thinkbot as an experimental AI thinking and reasoning crawler. The available primary evidence is a webmaster's report from August 2025. It says the bot may be related to thinkbot.agency but explicitly cannot confirm that connection. It also records that the bot did not look at robots.txt and asks site owners to block its IP addresses if the traffic causes trouble.

The report is valuable telemetry, not an operator statement. It does not establish that Thinkbot is currently active, that the 74 observed addresses remain associated with the bot, that the network blocks belong to an operator, or that the traffic is used for training. Do not reuse the report's site-specific Tencent network list as an allowlist or denylist without independent validation.

The observed header is:

configuration / code
Mozilla/5.0 (compatible; Thinkbot/0.5.8; +In_the_test_phase,_if_the_Thinkbot_brings_you_trouble,_please_block_its_IP_address._Thank_you.)

If your own logs confirm an exact Thinkbot token and you want to communicate a restriction, publish:

configuration / code
User-agent: Thinkbot
Disallow: /

Because the report says its observed client did not inspect robots.txt, treat this as communication rather than enforcement. Use authentication, authorization, rate limits, and edge controls for private or high-cost routes.

Layered verification

Start with your own raw access logs. Preserve the complete User-Agent, source IP, ASN, reverse DNS, method, path, status, response size, redirects, timestamp, and request rate. Compare the observed header byte-for-byte with the value in the report, but remember that a client can change or spoof it.

The report recorded traffic from many source addresses and network blocks. That is a signal to investigate operational impact, not proof of a stable operator range. Independently resolve and validate each address before taking network action. Do not conclude that every request from a named ASN is Thinkbot traffic, and do not treat a community-maintained list as current infrastructure evidence.

Analyze behavior without assigning purpose prematurely. Broad traversal, high concurrency, repeated retries, and requests for private or API paths may indicate scraping or abuse. A browser-like string and a reported test-phase message do not prove AI training, indexing, or an authorized experiment.

Evaluate /robots.txt independently even when the observed report says the bot did not. Confirm the canonical host, response status, content type, exact group, and path match. Page-level directives may express a discoverability preference:

configuration / code
<meta name="robots" content="noindex, nofollow">
configuration / code
X-Robots-Tag: noindex, nofollow

These directives do not protect private content and cannot substitute for active enforcement against a client that ignores robots rules.

WAF and Nginx remediation examples

If logs confirm unwanted requests carrying the exact token, a narrow WAF rule can block the self-declared identity:

configuration / code
{
  "description": "Block observed Thinkbot token",
  "expression": "lower(http.user_agent) contains \"thinkbot\"",
  "action": "block"
}

For Nginx, scope enforcement to private and high-cost routes while you investigate source networks:

configuration / code
map $http_user_agent $block_thinkbot {
    default 0;
    ~*Thinkbot 1;
}

server {
    location ~ ^/(private|internal|account|uploads|api)/ {
        if ($block_thinkbot) { return 403; }
        try_files $uri $uri/ =404;
    }
}

A User-Agent rule is easy to spoof or evade and may block an approved monitor that contains the substring. Avoid copying the report's broad IP blocks into production. If local logs show persistent abuse, validate individual addresses, apply rate limits or challenge controls, and review the effect on ordinary users before broad network blocking.

Review checklist

Search logs for Thinkbot and preserve the full header, source IP, ASN, requested paths, response sizes, status, and timing. Record whether the traffic is current, whether it reads robots.txt, and whether the behavior is broad or targeted. Treat the possible thinkbot.agency connection as unconfirmed and do not infer an operator from the header alone.

Decide whether your objective is to protect bandwidth, prevent possible content collection, secure private routes, or investigate experimental traffic. Publish a targeted robots group for communication, then enforce sensitive routes with WAF, Nginx, authentication, and rate limiting. Test public docs, media, feeds, sitemaps, uploads, and APIs separately. Revisit the profile if an operator publishes a current User-Agent, source verification, purpose statement, or opt-out policy.

References

  1. The Boston Diaries: “Bro, ban me at the IP level if you don't like me!” — first-hand webmaster observation of Thinkbot's header, robots behavior, and source-network spread; not an operator policy.
  2. Thinkbot Agency — possible association linked from the observation; attribution was not confirmed, so it is not used as evidence of crawler behavior.
  3. Google Robots.txt Introduction — general explanation of crawler directives and their limitations.

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.