PhxBot: Robots.txt & Crawl Policy Reference
Technical reference for the historical PhxBot registry label, with explicit limits around unavailable documentation and an unverified email-like User-Agent.
AI Summary: The registry describes PhxBot as a Phoenix web crawler for content discovery and records
PhxBot/0.1 (phxbot@protonmail.com), but a headless-browser search returned unrelated real-estate results and no first-party PhxBot source. The email-like suffix is not a verified contact. No current operator, robots policy, rate limit, IP range, reverse-DNS method, or opt-out process is established. Treat the token as a historical registry signal, not authenticated identity.
Role and policy boundary
The registry describes PhxBot as a Phoenix web crawler for content discovery and records this User-Agent:
PhxBot/0.1 (phxbot@protonmail.com)
The registry entry is the only role evidence available in this review. A headless Bing query for the exact version, email-like suffix, and crawler terms returned unrelated German real-estate search results and no apparent first-party PhxBot source. The email-like suffix was not treated as a verified contact, and the search result was not used as evidence of operation.
This profile therefore uses legacy-label. A matching request could be a legacy deployment, a test client, a private tool, or a spoofed header. Do not infer that a current Phoenix operator exists, that the client performs content discovery, or that it has any relationship to search indexing, AI input, model training, or data retention. A User-Agent is a label, not an authentication mechanism.
Because no current policy was verified, do not publish an assumed PhxBot robots group as if it came from an operator. If logs establish a current, attributable token and you decide to exclude it, use the exact observed value in a deliberate site-owner rule:
User-agent: CONFIRMED-PHXBOT
Disallow: /
For selective access after confirmation:
User-agent: CONFIRMED-PHXBOT
Allow: /public/
Allow: /docs/
Disallow: /admin/
Disallow: /account/
Disallow: /private/
Disallow: /licensed/
Disallow: /api/
These are explanatory site-owner controls, not recovered PhxBot instructions. Robots.txt is advisory and cannot secure private or licensed content; use authentication, authorization, signed URLs, data minimization, and origin controls.
Layered verification
Start with raw access logs and preserve the complete User-Agent, source IP, ASN, reverse-DNS result, HTTP method, requested path, response status, response size, redirect chain, timestamp, concurrency, and request rate. No current PhxBot source in this review published an IP range, reverse-DNS procedure, Crawl-delay, rate guidance, removal address, or verification workflow. The registry token is useful for triage but is not authentication.
Do not treat the Proton Mail address in the header as a verified operator contact. Check source IP and DNS evidence independently, and require a current authoritative response before creating a trust exception. An email-like string can be copied by unrelated clients and does not establish identity or permission.
Compare observed behavior with the historical content-discovery hypothesis without turning it into attribution. Public HTML, metadata, feeds, sitemaps, and ordinary assets may be requested by many tools. Private endpoints, licensed content, account routes, APIs, high concurrency, repeated retries, or unexpected bulk downloads establish impact and load risk, not Phoenix attribution or downstream use.
Evaluate /robots.txt only after the actual observed token is known. Confirm that it is served by the intended host, returns a successful text response, and contains the exact group you intend to publish. An absent operator source is not evidence of robots compliance. If no exact group exists, a global rule may affect unrelated clients and should be adopted only as an explicit site-wide decision.
Page-level directives can express indexing preferences:
<meta name="robots" content="noindex, nofollow">
X-Robots-Tag: noindex, nofollow
These signals do not authenticate an undocumented crawler or secure private paths. Enforce sensitive boundaries in the application and at the origin. If your policy distinguishes search indexing, AI input, reference use, and model training, document each purpose separately rather than inferring permission from a registry label.
WAF and Nginx remediation examples
When the source is unverified, use report-only logging and a narrow observation. Avoid broad phx, bot, email, or browser rules that can affect unrelated clients:
{
"description": "Observe unverified PhxBot candidates",
"expression": "lower(http.user_agent) contains \"phxbot/\"",
"action": "log"
}
After independent verification and a policy decision, scope enforcement to sensitive routes and preserve evidence for the rule:
map $http_user_agent $block_confirmed_phxbot_private {
default 0;
# Add only a complete, independently verified token here.
# ~*PhxBot/0\.1\s+\(phxbot@protonmail\.com\) 1;
}
server {
location ~ ^/(admin|account|private|licensed|internal|api)/ {
if ($block_confirmed_phxbot_private) { return 403; }
try_files $uri $uri/ =404;
}
}
A User-Agent match is easy to spoof, and the registry contact string is not verified. Do not invent an IP allowlist, reverse-DNS suffix, rate, training policy, or permanent trust exception. Use per-client rate limits, concurrency ceilings, timeouts, response-size controls, caching, and anomaly detection at the edge or origin. Start in report-only mode, review false positives, and restrict only after evidence supports the action.
Test public pages, feeds, sitemaps, structured data, licensed assets, account routes, APIs, 429 behavior, response-size limits, and approved integrations separately. Pair WAF controls with authentication and application authorization instead of using robots.txt as an access-control mechanism.
Review checklist
Search logs for the exact PhxBot/0.1 substring and preserve the complete header, source IP, ASN, PTR result, forward lookup, path, method, response size, status, timing, rate, concurrency, and redirects. Treat the email-like suffix as an unverified string, not a support channel.
Re-check the crawler registry, any future Phoenix/OpenHose source, robots.txt, and operator contact path. In this review no relevant first-party source was found and the headless search returned unrelated results. Keep this profile at legacy-label until a current source publishes an exact token, purpose, robots behavior, source-verification method, rate guidance, or opt-out process.
Decide whether your objective is to preserve public discovery, reduce crawl load, limit extraction, protect licensed material, or prevent private access. Publish exact robots rules only after the token is known; enforce private routes with authentication and origin controls.
Review search indexing, AI input, reference use, and model-training decisions separately. Neither the registry description nor an email-like header establishes downstream permission or use. Do not claim successful blocking or verification from configuration alone; validate later logs and obtain a current operator response when possible.
Record the date and reason for the legacy-label classification so future evidence can be compared without silently upgrading an unsupported historical label. If a new deployment identifies itself differently, create a separate evidence trail.
References
- Crawler User Agents community registry — registry context for the PhxBot label and User-Agent; it does not authenticate current infrastructure.
- Headless Bing search for PhxBot — reviewed on 2026-08-25; results were unrelated and were not used as proof of operation.
- Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
- RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.
Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.