StractBot: Robots.txt & Crawl Policy Reference
Technical reference for StractBot, a web crawler used for a PR tool, including its robots.txt compliance and the transition from its historical domain.
AI Summary:
StractBotis a web crawler operated by Stract. While the registry describes it as an open-source search engine, the current official documentation states it indexes public content for a PR tool for companies. The operator confirms it respects theStractBotrobots.txt token. Because the operator does not publish a static IP list, reverse-DNS method, or exact full User-Agent string, verify the source network independently before trusting the header.
Role and policy boundary
The registry describes this entry as a Stract search engine web crawler and records this historical User-Agent:
Mozilla/5.0 (compatible; StractBot/0.1; open source search engine; +https://trystract.com/webmasters)
During headless-browser review, the registry-linked trystract.com domain returned an SSL certificate error. The current official documentation is located at https://stract.com/bot. This active page clarifies that StractBot indexes public content to power a PR tool for companies, rather than operating as a general-purpose search engine.
The operator explicitly states that StractBot respects the standard robots.txt protocol and identifies itself using the StractBot token. For selective access to public pages:
User-agent: StractBot
Allow: /public/
Allow: /press/
Disallow: /admin/
Disallow: /account/
Disallow: /private/
Disallow: /api/
For a complete exclusion from the PR tool index:
User-agent: StractBot
Disallow: /
These are site-owner controls. Robots.txt is advisory and cannot secure private or licensed content; use authentication, authorization, signed URLs, data minimization, and origin controls for those boundaries. The business-intelligence purpose does not establish permission for generative AI answers or third-party model training.
Layered verification
Preserve the complete User-Agent, source IP, ASN, reverse-DNS result, HTTP method, requested path, response status, response size, redirect chain, timestamp, concurrency, and request rate in logs. The reviewed Stract documentation does not publish a static IP list, a specific reverse-DNS verification procedure, or rate guidance.
Because a User-Agent is easily spoofed, do not treat the header alone as authentication. Check source IP and DNS evidence independently. A StractBot header from a residential network, a known proxy, or a cloud provider without an established Stract affiliation should not be treated as authentic traffic.
Compare observed traffic with the documented PR-tool indexing role. Public HTML, press releases, metadata, feeds, and ordinary assets are consistent with discovery. Private endpoints, account routes, APIs, licensed content, high concurrency, repeated retries, or unexpected bulk downloads establish impact and load risk, not permission.
Evaluate /robots.txt independently. Confirm that your host returns a successful text response and contains the exact StractBot group you intend to publish. Check precedence, path matching, and actual request behavior. Meta directives can express indexing preferences:
<meta name="robots" content="noindex, nofollow">
X-Robots-Tag: noindex, nofollow
These signals do not authenticate the crawler or protect private routes. Enforce sensitive boundaries in the application and at the origin.
WAF and Nginx remediation examples
Begin in report-only mode and correlate the StractBot User-Agent with independent network verification. Adapt the expression to your WAF provider:
{
"description": "Observe StractBot candidates before enforcement",
"expression": "lower(http.user_agent) contains \"stractbot\"",
"action": "log"
}
For a deliberately restricted private route, use the exact token only after verifying the source network:
map $http_user_agent $block_stract_private {
default 0;
~*StractBot 1;
}
server {
location ~ ^/(admin|account|private|licensed|internal|api)/ {
if ($block_stract_private) { return 403; }
try_files $uri $uri/ =404;
}
}
This rule is route-scoped and does not prove identity. Do not invent an IP allowlist, Stract ASN exception, rate, or training permission without operator documentation.
A User-Agent is easy to spoof. Use per-client rate limits, concurrency ceilings, timeouts, response-size controls, caching, and anomaly detection at the edge or origin. Start in report-only mode, review false positives, and restrict only after evidence supports the action.
Test public pages, feeds, sitemaps, structured data, licensed assets, account routes, APIs, 429 behavior, response-size limits, and approved integrations separately. Pair WAF controls with authentication and application authorization instead of using robots.txt as an access-control mechanism.
Review checklist
Search logs for the exact StractBot string, preserving the full header, source IP, ASN, PTR result, forward lookup, path, method, response size, status, timing, rate, concurrency, and redirects. Treat the historical trystract.com URL in the registry header as outdated.
Verify the source network independently before trusting the client. Treat requests from unverified networks as spoofed.
Check your robots file for an exact StractBot group. Test precedence and path behavior, and verify actual logs after publishing. Use META or X-Robots-Tag for indexing preferences and authentication for private resources.
Review business-intelligence indexing, AI input, reference use, and model-training decisions separately. The operator explicitly states the crawler is used for a PR tool for companies.
Keep the profile at documented-limit: the operator publishes the crawler identity, robots compliance, and purpose, but lacks a static IP list, exact full UA pattern, or reverse-DNS verification method. Re-check the crawler documentation when the source network or traffic pattern changes. Do not claim successful blocking or verification from configuration alone; validate later logs.
References
- Stract Crawler Information — current official page documenting the StractBot robots.txt token and PR tool purpose; reviewed with the headless browser on 2026-08-25.
- Crawler User Agents community registry — registry context for the StractBot label and historical User-Agent.
- Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
- RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.
Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.