OrangeBot: Robots.txt & Crawl Policy Reference
Technical reference for the OrangeBot registry label, separating Orange’s current corporate site and global robots rules from an unverified search-crawler identity.
AI Summary: The registry labels OrangeBot as an Orange search-engine crawler and records an incomplete browser-compatible token with an email-like suffix, but current Orange corporate pages reviewed here do not document that bot identity. Orange’s robots.txt has a global
User-agent: *group, many path rules, an Algolia verification comment, and a sitemap, but no OrangeBot group, Crawl-delay, IP range, or reverse-DNS method. Treat the registry token as a candidate signal and verify observed traffic independently.
Role and policy boundary
The registry describes OrangeBot as an Orange search-engine web crawler and records this value:
Mozilla/5.0 (compatible; OrangeBot/2.0; support.orangebot@orange.com
The value is syntactically incomplete and is the only bot-specific identity in the registry entry. The current Orange corporate site confirms Orange as a telecommunications and digital-services group, but the reviewed pages do not publish an OrangeBot specification or confirm that this exact header is current. Do not silently correct the missing closing character, add a URL, or treat the email-like suffix as a verified contact.
The current https://www.orange.com/robots.txt contains a global group with asset allows and many path disallows, including administration, search, login, user, research, and other application paths. It also includes an Algolia-Crawler-Verif comment and a sitemap. Those are site-specific controls and a verification comment for another crawler context; they are not proof that OrangeBot is operated by Orange or that OrangeBot obeys them.
A matching request could be a legacy deployment, a test client, an Orange internal tool, an unrelated browser-like client, or a spoofed header. Do not merge it with Orange’s corporate website traffic, Algolia crawler activity, or other telecom/search services. The token does not establish search indexing, AI input, model training, data retention, or permission to access private customer or corporate content.
Because no current OrangeBot policy was verified, do not publish an assumed bot-specific rule as if it came from Orange. If logs establish a current, attributable token and you decide to exclude it, use the exact observed value in a deliberate site-owner rule:
User-agent: CONFIRMED-ORANGEBOT
Disallow: /
For selective access after confirmation:
User-agent: CONFIRMED-ORANGEBOT
Allow: /public/
Allow: /docs/
Disallow: /admin/
Disallow: /account/
Disallow: /private/
Disallow: /licensed/
Disallow: /api/
These are explanatory site-owner controls, not recovered Orange instructions. Robots.txt is advisory and cannot secure private or licensed content; use authentication, authorization, signed URLs, data minimization, and origin controls.
Layered verification
Start with raw access logs and preserve the complete User-Agent, source IP, ASN, reverse-DNS result, HTTP method, requested path, response status, response size, redirect chain, timestamp, concurrency, and request rate. No current Orange source in this review published an OrangeBot IP range, forward-confirmed reverse-DNS procedure, Crawl-delay, rate guidance, or removal workflow. The registry token is useful for triage but is not authentication.
Do not treat orange.com branding or an email-like suffix in a header as proof of Orange ownership. Check source IP and DNS evidence independently, and require a current first-party verification path before creating a trust exception. An IP associated with Orange is not alone proof that a request is an authorized search crawler.
Compare observed behavior with the claimed search-crawling hypothesis without turning it into attribution. Public HTML, metadata, feeds, sitemaps, and static assets may be requested by many clients. Private customer routes, account endpoints, APIs, licensed assets, high concurrency, repeated retries, or unexpected bulk downloads establish impact and load risk, not Orange identity or downstream use.
Evaluate /robots.txt independently. Confirm that your intended host returns a successful text response and contains the exact OrangeBot group you intend to publish. The reviewed Orange file has a global group and many site-specific rules, but no OrangeBot group. The Algolia verification comment must not be reused as an OrangeBot authentication mechanism.
Page-level directives can express indexing preferences:
<meta name="robots" content="noindex, nofollow">
X-Robots-Tag: noindex, nofollow
These signals do not authenticate an undocumented crawler or secure private customer paths. Enforce sensitive boundaries in the application and at the origin. If your policy distinguishes search indexing, AI input, reference use, and model training, document each purpose separately rather than inferring permission from Orange branding or the registry description.
WAF and Nginx remediation examples
When the identity is unverified, use report-only logging and a narrow observation. Avoid broad orange, Mozilla, email, or telecom rules that can affect unrelated clients:
{
"description": "Observe unverified OrangeBot candidates",
"expression": "lower(http.user_agent) contains \"orangebot\"",
"action": "log"
}
After independent verification and a policy decision, scope enforcement to sensitive routes and preserve the evidence supporting the rule:
map $http_user_agent $block_confirmed_orangebot_private {
default 0;
# Add only a complete, independently verified token here.
# ~*OrangeBot/2\.0 1;
}
server {
location ~ ^/(admin|account|private|licensed|internal|api)/ {
if ($block_confirmed_orangebot_private) { return 403; }
try_files $uri $uri/ =404;
}
}
A User-Agent match is easy to spoof, and the registry value is incomplete. Do not invent an Orange IP allowlist, ASN exception, reverse-DNS suffix, rate, training policy, or permanent trust rule. Use per-client rate limits, concurrency ceilings, timeouts, response-size controls, caching, and anomaly detection at the edge or origin. Start in report-only mode, review false positives, and restrict only after evidence supports the action.
Test public pages, feeds, sitemaps, structured data, licensed assets, customer routes, account paths, APIs, 429 behavior, response-size limits, and approved integrations separately. Pair WAF controls with authentication and application authorization instead of using robots.txt as an access-control mechanism.
Review checklist
Search logs for the exact OrangeBot substring and preserve the complete header, source IP, ASN, PTR result, forward lookup, path, method, response size, status, timing, rate, concurrency, and redirects. Treat the incomplete registry header and corporate branding as clues, not proof of ownership.
Re-check Orange’s current crawler documentation, the registry source, robots.txt, and any future operator contact path. In this review the corporate website was accessible but did not document OrangeBot; its robots file had no OrangeBot-specific group. Keep the profile at documented-limit until a current source establishes the token, purpose, robots behavior, source-verification method, rate guidance, or opt-out process.
Check your robots file for an exact OrangeBot group, test precedence and path behavior, and measure origin load. Do not generalize Orange’s global rules or its Algolia verification comment to third-party sites. Use META or X-Robots-Tag for indexing preferences and authentication for private content.
Review search indexing, AI input, reference use, and model-training decisions separately. Neither the registry description nor the current corporate page establishes downstream permission or use. Do not claim successful blocking or verification from configuration alone; validate later logs and obtain a current operator response when possible.
Decide whether your objective is to preserve public discovery, reduce crawl load, limit extraction, protect licensed material, or prevent private access. Publish exact robots rules only after the token is known and enforce sensitive routes at the application and origin layers.
References
- Orange corporate website — current first-party corporate context; reviewed with the headless browser on 2026-08-25, without OrangeBot documentation.
- Orange robots.txt — current global rules, path restrictions, Algolia verification comment, and sitemap; no OrangeBot-specific group, reviewed with the headless browser on 2026-08-25.
- Crawler User Agents community registry — registry context for the incomplete OrangeBot token; it does not authenticate current Orange infrastructure.
- Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
- RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.
Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.