MagiBot: Robots.txt & Crawl Policy Reference
Technical reference for the historical MagiBot registry label, with explicit limits around unavailable Peak Labs documentation and current request verification.
AI Summary:
MagiBotis a historical registry label associated with Peak Labs information extraction, but its linked documentation timed out and thewwwvariant closed the connection during review. No current operator policy, robots behavior, IP range, or source-verification method was established. Treat the registered User-Agent as an unverified historical signal and protect sensitive data with application controls.
Role and policy boundary
The inventory describes MagiBot as a Peak Labs information-extraction crawler and records this User-Agent:
Mozilla/5.0 (compatible; MagiBot; +https://www.magi.com/)
The linked https://magi.com/bots documentation request timed out after 90 seconds in the headless browser. The https://www.magi.com/bots variant then failed with net::ERR_CONNECTION_CLOSED. No current operator policy, canonical bot page, source network, robots instructions, rate guidance, or opt-out process could be read.
Information extraction is therefore a registry description, not a verified current purpose. A request carrying the token may come from a legacy deployment, a data integration, a test client, a fork, or a spoofed header. Do not infer current activity, operator identity, downstream use, AI-training purpose, permission to crawl, or access to private content from the token or hostname.
If access logs confirm the exact token and your policy is to exclude it from discovery, a narrow defensive group could be:
User-agent: MagiBot
Disallow: /
For selective access to approved public pages while protecting sensitive data:
User-agent: MagiBot
Allow: /public/
Allow: /docs/
Disallow: /admin/
Disallow: /account/
Disallow: /private/
Disallow: /licensed/
Disallow: /api/
These are site-owner examples, not recovered MagiBot instructions. Robots.txt is advisory and cannot protect private or licensed content; use authentication, authorization, signed URLs, data minimization, and origin controls for those boundaries.
Layered verification
Begin with raw access logs and preserve the complete User-Agent, source IP, ASN, reverse-DNS result, HTTP method, requested path, response status, response size, redirect chain, timestamp, and request rate. No reachable current MagiBot source published an IP range, reverse-DNS procedure, rate policy, or canonical contact. The registry header is useful for triage but is not authentication.
Compare observed behavior with an information-extraction hypothesis without turning it into attribution. Public HTML, structured data, feeds, and ordinary assets may be requested by many tools. Private data exports, licensed documents, account routes, APIs, high concurrency, repeated retries, or unexpected bulk downloads may indicate unauthorized extraction, spoofing, abuse, or a different client. These observations establish data risk and impact, not operator identity or downstream use.
Evaluate /robots.txt independently if you publish a rule. Confirm that it is served from the intended host, returns a successful status and text content type, and contains the exact group you intend to apply. Test the complete token and global group separately. Page-level directives can express discovery preferences:
<meta name="robots" content="noindex, nofollow">
X-Robots-Tag: noindex, nofollow
These signals do not authenticate MagiBot or secure private routes. Enforce sensitive boundaries in the application and at the origin, and avoid placing secret data in public HTML merely because a robots rule exists.
WAF and Nginx remediation examples
If logs show the historical token, start with report-only mode and correlate it with source evidence and data sensitivity. Adapt the expression to your WAF provider:
{
"description": "Review historical MagiBot extraction candidates",
"expression": "lower(http.user_agent) contains \"magibot\"",
"action": "log"
}
After reviewing false positives and confirming the business decision, scope enforcement to sensitive routes:
map $http_user_agent $block_magibot_private {
default 0;
~*MagiBot 1;
}
server {
location ~ ^/(admin|account|private|licensed|internal|api)/ {
if ($block_magibot_private) { return 403; }
try_files $uri $uri/ =404;
}
}
A User-Agent match is easy to spoof and may catch an authorized Peak Labs integration or a local test. Do not invent an IP allowlist, reverse-DNS suffix, or current operator exception from a timeout. Test public HTML, structured data, feeds, sitemaps, licensed documents, accounts, APIs, and approved integrations separately. Pair edge controls with authentication, rate limits, signed assets, caching, data-loss monitoring, and anomaly detection.
Review checklist
Search logs for the complete registered header and preserve representative source IPs, ASNs, reverse-DNS results, paths, methods, response sizes, statuses, timing, and rate. Check whether requested content is public and intended for extraction, and whether the source can be independently verified. Do not classify a request as authentic from the token alone.
Re-check magi.com, www.magi.com, the /bots path, and the community registry when a credible operator source becomes reachable. Keep this profile at legacy-label until a first-party source publishes a current policy, canonical User-Agent, source-verification method, IP ranges, rate guidance, or opt-out process. Treat the timeout and connection closure as source-access limitations, not proof that the crawler is inactive.
Decide whether your objective is to preserve public discovery, limit extraction, protect licensed content, or reduce crawl load. Publish exact robots rules and enforce private routes with application and WAF controls. Do not claim successful blocking from configuration alone; verify subsequent access logs and response behavior.
The registry’s information-extraction label does not establish AI-model training use. Keep any downstream-use statement separate from the limited evidence available here.
References
- MagiBot documentation — registry-linked source; headless-browser request timed out after 90 seconds on 2026-08-25.
- MagiBot documentation www variant — direct headless-browser request failed with
net::ERR_CONNECTION_CLOSEDon 2026-08-25. - Crawler User Agents community registry — registry context for the label and User-Agent; it does not authenticate current Magi infrastructure.
- Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
- RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.
Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.