AddSearchBot: Robots.txt & Crawl Policy Reference
Technical reference for AddSearchBot. Learn how to investigate site search crawling traffic when the AddSearch documentation is partially inaccessible.
AI Summary:
AddSearchBotis the crawler associated with AddSearch, a site search solution. The registry-linked policy URL redirects to a whitelisting documentation page, but the full content was inaccessible during review. While historically identifying asMozilla/5.0 (compatible; AddSearchBot/0.9; +http://www.addsearch.com/bot; info@addsearch.com), its current IP ranges and exact robots behavior could not be fully verified. Treat any local observation as evidence to investigate and use layered controls.
Role and policy boundary
The registry describes AddSearchBot as a web crawler for AddSearch site search indexing. While the domain addsearch.com exists and redirects to a whitelisting documentation page (https://www.addsearch.com/docs/indexing/whitelisting-addsearch-bot/), the full content of that page was inaccessible (returned blank or blocked by CAPTCHA) during headless-browser review.
No canonical source IP list, crawl schedule, rate limit, robots statement, or verification method was confirmed from the primary source. If you use AddSearch for your website, this crawler is likely authorized traffic indexing your content for your own search functionality. Do not infer that it collects AI-training data or builds a public search index.
If your logs confirm an exact AddSearchBot token and you want to communicate a restriction, publish:
User-agent: AddSearchBot
Disallow: /
For selective access (e.g., to index only public pages for your site search):
User-agent: AddSearchBot
Allow: /public-reference/
Allow: /docs/
Disallow: /private/
Disallow: /internal/
Disallow: /api/
Robots.txt is advisory and cannot protect private or licensed content. Use authentication, authorization, signed URLs, and origin controls for those boundaries.
Layered verification
Start with raw access logs and preserve the complete User-Agent, source IP, ASN, reverse DNS, method, path, status, response size, redirects, timestamp, and request rate. Because the primary documentation was inaccessible, do not classify traffic from a partial substring alone.
Analyze behavior without assigning purpose prematurely. Requests for public pages, feeds, sitemaps, and metadata may resemble authorized site search indexing; deep traversal of private areas, high concurrency, repeated retries, original-asset downloads, or private API access may indicate misconfiguration or abuse. These patterns demonstrate operational impact but cannot prove the operator or downstream use.
Evaluate /robots.txt independently. Confirm the canonical host, response status, content type, exact user-agent group, and path match. A page-level directive may express a discoverability preference:
<meta name="robots" content="noindex, nofollow">
X-Robots-Tag: noindex, nofollow
These signals do not establish an opt-out for an undocumented client and do not secure private routes. Use authenticated delivery, signed URLs, and application authorization.
WAF and Nginx remediation examples
Once logs confirm an exact unwanted token, a narrow WAF rule can block the declared identity. Replace the example expression if your observed header differs:
{
"description": "Block observed AddSearchBot token",
"expression": "lower(http.user_agent) contains \"addsearchbot\"",
"action": "block"
}
For Nginx, scope enforcement to private and high-cost routes while investigating public access:
map $http_user_agent $block_addsearchbot {
default 0;
~*AddSearchBot 1;
}
server {
location ~ ^/(private|internal|account|uploads|paywall|api)/ {
if ($block_addsearchbot) { return 403; }
try_files $uri $uri/ =404;
}
}
A User-Agent rule is easy to spoof or evade and may block a legitimate client using the substring. Do not create an IP allowlist or broad network block without current operator evidence. Test browsers, social previews, feed readers, search crawlers, approved monitors, and customer integrations. Pair edge matching with authentication, rate limits, signed assets, and anomaly detection.
Review checklist
Search logs for every exact header that may be associated with AddSearchBot and preserve representative requests. Record paths, response sizes, statuses, source networks, timing, and rate. Re-check the registry-linked policy URL before upgrading this profile; during this review the documentation page was inaccessible.
Decide whether your objective is to preserve authorized site search indexing, prevent extraction, protect private content, or reduce crawl load. Publish a targeted robots group only for the exact observed token, enforce sensitive routes with WAF and application controls, and test docs, media, feeds, sitemaps, uploads, and APIs separately. Revisit the profile if AddSearch publishes an accessible live policy, canonical User-Agent, source verification method, purpose statement, or robots behavior.
References
- Registry-linked AddSearch policy URL — redirected to a documentation page that was inaccessible during the 2026-08-24 headless-browser review.
- Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
- RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.
Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.