HenkBot: Robots.txt & Crawl Policy Reference
Technical reference for HenkBot, Valyu's web indexing crawler. Learn its published User-Agent, robots.txt behavior, and safe server controls.
AI Summary: HenkBot is Valyu's documented web crawler for powering its web indexing and real-time search APIs. Valyu publishes the exact User-Agent, states that HenkBot always honors
robots.txt, respectsCrawl-delay, and followsnoindexandnofollowdirectives. Use a targeted robots rule to allow or block it, and verify the declared identity in logs before adding edge enforcement.
Role and policy boundary
Valyu describes HenkBot as the crawler behind its web indexing system. It crawls and indexes content so that Valyu's APIs can return current search results to AI applications. This role is different from a one-off browser fetch: HenkBot is an indexing crawler, so its requests may recur as Valyu refreshes its index.
The official Valyu documentation publishes both the exact User-Agent and the policy controls it recognizes. It states that HenkBot uses adaptive rate limiting, honors robots.txt, respects a specified Crawl-delay, and follows noindex and nofollow metadata. These are documented service behaviors, not assumptions derived from a third-party bot list.
A site owner can therefore make a deliberate search-visibility choice. Allowing HenkBot may make public content available to applications using Valyu's search APIs. Blocking it removes that source from Valyu's index but does not block other search engines, other AI crawlers, or direct user requests.
To allow the crawler across the public site:
User-agent: HenkBot
Allow: /
To block it everywhere:
User-agent: HenkBot
Disallow: /
For a selective policy, keep public documentation available while excluding sensitive areas:
User-agent: HenkBot
Allow: /docs/
Allow: /public-reference/
Disallow: /admin/
Disallow: /account/
Disallow: /api/
Layered verification
Begin with access logs and match the complete published token, including the +https://valyu.ai/crawler contact URL where present. A User-Agent is self-declared, so it is a useful classification signal but not proof of origin. Record source IP, reverse DNS, request path, status, response size, timestamp, and request rate before treating a request as verified.
Next, fetch the canonical /robots.txt as the site would serve it to a crawler. Check that the HenkBot group is syntactically valid, that the requested path falls under the intended rule, and that the response is available at the canonical host. Valyu's stated Crawl-delay behavior is useful only if the rule is published in a way the crawler can read and the crawler is actually using the declared identity.
Inspect page-level directives as a separate layer. A page that should not be indexed can use:
<meta name="robots" content="noindex, nofollow">
or:
X-Robots-Tag: noindex, nofollow
These signals are not substitutes for authentication. If a route contains private or high-cost data, enforce access with authorization and rate limits even when HenkBot is documented as robots-compliant.
WAF and Nginx remediation examples
If logs confirm unwanted traffic that still declares HenkBot, a narrow WAF rule can provide active enforcement. Use it only after checking whether you intended to block Valyu search visibility:
{
"description": "Block HenkBot at the edge",
"expression": "lower(http.user_agent) contains \"henkbot\"",
"action": "block"
}
For Nginx, match the token without depending on the full browser-compatible prefix:
map $http_user_agent $block_henkbot {
default 0;
~*HenkBot 1;
}
server {
location / {
if ($block_henkbot) { return 403; }
try_files $uri $uri/ =404;
}
}
Test the rule on a staging host first. A User-Agent match can be spoofed and can also block a legitimate Valyu request if you later decide to allow indexing. For higher-confidence attribution, combine log review and source verification with the header rule rather than presenting the header alone as authentication.
Review checklist
Confirm that your policy decision is intentional: allow HenkBot for Valyu search visibility, block it for all routes, or allow only public documentation. Publish the exact User-agent: HenkBot group in the canonical /robots.txt, and add Crawl-delay only where it matches your server's operational needs.
After publishing, request representative public and private paths and compare the result with your expected directives. Review logs for the exact User-Agent, rate-limit behavior, response status, and source network. If you need active blocking, deploy the narrow WAF or Nginx rule and re-test after CDN changes. Keep search, training, and direct assistant policies separate: blocking HenkBot changes Valyu indexing, not every form of AI access.
References
- Valyu Crawler Documentation — official User-Agent, robots.txt examples, crawl behavior, and crawler contact guidance.
- Google Robots.txt Introduction — general explanation of crawler access directives and their limits.
Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.