ShapBot: Robots.txt & Crawl Policy Reference
Technical reference for ShapBot, Parallel's documented crawler for discovering and indexing websites for its web APIs. Learn the User-Agent, IP verification, and robots controls.
AI Summary: ShapBot is Parallel's documented crawler for discovering and indexing websites used by Parallel's web APIs. Parallel publishes a stable
ShapBottoken, the full versioned User-Agent, a changing IP list, and support contact. It recommends permitting ShapBot for search visibility and says webmasters should verify both the User-Agent and designated source IP before allowlisting.
Role and policy boundary
Parallel's official webmaster documentation describes ShapBot as the crawler that helps discover and index websites for Parallel's web APIs. The stated benefit of allowing it is visibility in Parallel search results. This is an indexing and web-API context, not a claim that ShapBot is used to train a foundation model.
The stable token is ShapBot, while the documented full example is:
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; ShapBot/0.1.0
Parallel recommends allowing ShapBot when a site wants to maximize visibility in its search results. To allow public documentation while excluding private and operational paths:
User-agent: ShapBot
Allow: /docs/
Allow: /public-reference/
Disallow: /private/
Disallow: /internal/
Disallow: /account/
Disallow: /api/
To block the crawler entirely:
User-agent: ShapBot
Disallow: /
The robots decision is separate from the network decision. Parallel's documentation says that permitting a request through a firewall or WAF does not override a disallow rule. Robots.txt is advisory and cannot protect confidential content; use authentication and authorization for that boundary.
Layered verification
Parallel publishes ShapBot source addresses at shapbot.json. The documentation recommends verifying both conditions: the User-Agent contains the ShapBot token and the source IP belongs to the designated current ranges. Because the IP list can change, retrieve it through a controlled allowlist update process rather than treating a copied list as permanent.
Record the full User-Agent, source IP, ASN, reverse DNS, request path, method, status, response size, redirect chain, timestamp, and request rate. A header alone is self-declared and can be spoofed. An IP alone may be shared or changed, so the two signals should be evaluated together.
Evaluate /robots.txt independently on each host. Confirm status, content type, canonical host/protocol/port, exact User-agent: ShapBot group, and path matching. For page-level discoverability, inspect:
<meta name="robots" content="noindex, nofollow">
X-Robots-Tag: noindex, nofollow
These directives do not replace the crawler-specific robots policy and do not secure private routes.
WAF and Nginx remediation examples
If you want to block Parallel's indexer, use the stable token rather than matching the complete versioned header:
{
"description": "Block Parallel ShapBot crawler",
"expression": "lower(http.user_agent) contains \"shapbot\"",
"action": "block"
}
For Nginx, scope a temporary block to high-cost paths while reviewing the business impact:
map $http_user_agent $is_shapbot {
default 0;
~*ShapBot 1;
}
server {
location ~ ^/(private|internal|account|uploads|api)/ {
if ($is_shapbot) { return 403; }
try_files $uri $uri/ =404;
}
}
If the goal is to allow verified ShapBot traffic, maintain the source-IP condition from the current shapbot.json file and still ensure robots.txt permits the path. Do not hard-code an IP list copied on one date. Test ordinary browsers, search crawlers, social previews, and approved integrations. For crawler questions or suspected misbehavior, Parallel lists support@parallel.ai as its support address.
Review checklist
Decide whether visibility in Parallel's web APIs and search results is valuable for your public content. Publish a host-specific ShapBot group that allows selected paths or blocks the crawler, and remember that a firewall allowlist cannot override robots.txt.
Verify observed traffic with both the stable ShapBot token and the current source ranges in shapbot.json. Record the full request, status, response size, path, and timing. Do not match a versioned full header because it may change. Protect private pages, uploads, APIs, and customer data with authentication regardless of crawler policy.
Re-test allowed and excluded paths after the robots file has propagated, verify subdomains separately, and monitor changes to Parallel's documentation. Contact support@parallel.ai with timestamps, source IPs, full User-Agent, and logs when investigating unexpected traffic.
References
- Parallel Crawler Documentation — official ShapBot purpose, User-Agent, robots guidance, IP-list link, and support contact.
- ShapBot IP ranges — official source list for network verification; refresh it because the documentation says crawler behavior and policies may change.
- RFC 9309 — Robots Exclusion Protocol standard relevant to robots.txt interpretation.
Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.