AspiegelBot: Robots.txt & Crawl Policy Reference
Technical reference for AspiegelBot, a web crawler operated by Aspiegel (Huawei). Learn how to manage its access for services like Petal Search.
AI Summary:
AspiegelBotis a web crawler operated by Aspiegel SE, a Huawei subsidiary that provides digital services in Europe, including Petal Search. While the corporate site (aspiegel.com) exists, specific crawler documentation is not published there. It historically identifies asMozilla/5.0 ... (compatible; AspiegelBot). Treat it as search engine traffic (similar to PetalBot), but use layered controls since exact IP ranges are not publicly documented on the linked site.
Role and policy boundary
The registry links AspiegelBot to Aspiegel SE, the European provider for Huawei Mobile Services. Headless-browser review confirmed that aspiegel.com is the official corporate site, but it does not host a dedicated crawler policy or IP list.
Given Aspiegel's role in operating Petal Search and other Huawei digital services, this bot is likely used for content discovery and search indexing. Blocking it may reduce your visibility in Huawei's search ecosystem. Do not infer that it collects AI-training data without further evidence.
If your logs confirm an exact AspiegelBot token and you want to communicate a restriction, publish:
User-agent: AspiegelBot
Disallow: /
For selective access (e.g., allowing it only on public content):
User-agent: AspiegelBot
Allow: /public-reference/
Allow: /articles/
Disallow: /private/
Disallow: /internal/
Disallow: /api/
Robots.txt is advisory and cannot protect private or licensed content. Use authentication, authorization, signed URLs, and origin controls for those boundaries.
Layered verification
Start with raw access logs and preserve the complete User-Agent, source IP, ASN, reverse DNS, method, path, status, response size, redirects, timestamp, and request rate. Because the primary documentation lacks an IP list, you cannot rely solely on published ranges. Look for reverse DNS entries associated with Huawei or Aspiegel (e.g., aspiegel.com or huawei.com).
Analyze behavior without assigning purpose prematurely. Requests for public pages, feeds, sitemaps, and metadata may resemble authorized search indexing; deep traversal of private areas, high concurrency, repeated retries, original-asset downloads, or private API access may indicate misconfiguration or spoofing. These patterns demonstrate operational impact but cannot prove the operator or downstream use.
Evaluate /robots.txt independently. Confirm the canonical host, response status, content type, exact user-agent group, and path match. A page-level directive may express a discoverability preference:
<meta name="robots" content="noindex, nofollow">
X-Robots-Tag: noindex, nofollow
These signals do not establish an opt-out for an undocumented client and do not secure private routes. Use authenticated delivery, signed URLs, and application authorization.
WAF and Nginx remediation examples
Once logs confirm an exact unwanted token (e.g., a spoofed bot or if you do not want to be indexed in Huawei services), a narrow WAF rule can block the declared identity. Replace the example expression if your observed header differs:
{
"description": "Block observed AspiegelBot token",
"expression": "lower(http.user_agent) contains \"aspiegelbot\"",
"action": "block"
}
For Nginx, scope enforcement to private and high-cost routes while investigating public access:
map $http_user_agent $block_aspiegelbot {
default 0;
~*AspiegelBot 1;
}
server {
location ~ ^/(private|internal|account|uploads|paywall|api)/ {
if ($block_aspiegelbot) { return 403; }
try_files $uri $uri/ =404;
}
}
A User-Agent rule is easy to spoof or evade and may block a legitimate client using the substring. Do not create an IP allowlist or broad network block without current operator evidence. Test browsers, social previews, feed readers, search crawlers, approved monitors, and customer integrations. Pair edge matching with authentication, rate limits, signed assets, and anomaly detection.
Review checklist
Search logs for every exact header that may be associated with AspiegelBot and preserve representative requests. Record paths, response sizes, statuses, source networks, timing, and rate. Re-check the registry-linked corporate URL for updated documentation; during this review, no specific crawler policy was found.
Decide whether your objective is to preserve search visibility in Huawei ecosystems, prevent extraction, protect private content, or reduce crawl load. Publish a targeted robots group only for the exact observed token, enforce sensitive routes with WAF and application controls, and test docs, media, feeds, sitemaps, uploads, and APIs separately. Revisit the profile if Aspiegel publishes an accessible live policy, canonical User-Agent, source verification method, purpose statement, or robots behavior.
References
- Aspiegel Corporate Site — verified as the official corporate site during the 2026-08-24 headless-browser review, though lacking specific crawler documentation.
- Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
- RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.
Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.