IstellaBot: Robots.txt & Crawl Policy Reference
Technical reference for the historical IstellaBot search-indexing label, with explicit limits around unavailable Tiscali documentation and current request verification.
AI Summary:
IstellaBot/1.23.15is a historical registry User-Agent associated with Istella search indexing. The registry-linked Tiscali source was blocked by browser policy, and a headless Google search was captcha-blocked, so no current operator policy, robots behavior, IP range, or verification method was established. Treat matching traffic as an unverified historical observation.
Role and policy boundary
The inventory associates IstellaBot with Istella web search and records this historical header:
Mozilla/5.0 (compatible; IstellaBot/1.23.15 +http://www.tiscali.it/)
The registry points to the Tiscali homepage rather than a bot-specific policy. Direct headless-browser access to that URL was blocked by policy restrictions during review. A follow-up headless Google search for an official IstellaBot robots page returned a captcha, so no search result or third-party claim was treated as evidence. These are source-access limitations, not proof that the historical client never existed or that all matching traffic is inactive.
Treat search indexing as the registry’s historical role hypothesis only. A matching header may come from an old deployment, a partner integration, a test client, a fork, or a spoofed request. Do not infer current activity, operator identity, downstream use, AI-training purpose, permission to crawl, or access to private content from the token.
If access logs confirm this exact token and your policy is to exclude it from discovery, a narrow defensive group could be:
User-agent: IstellaBot
Disallow: /
For selective access to approved public pages:
User-agent: IstellaBot
Allow: /public/
Allow: /docs/
Disallow: /admin/
Disallow: /account/
Disallow: /private/
Disallow: /api/
These are site-owner examples, not recovered Istella instructions. Robots.txt is advisory and cannot protect private or licensed content; use authentication, authorization, signed URLs, and origin controls for those boundaries.
Layered verification
Begin with raw access logs and preserve the complete User-Agent, source IP, ASN, reverse-DNS result, HTTP method, requested path, response status, response size, redirect chain, timestamp, and request rate. No current Istella or Tiscali source reviewed for this profile published a network range, reverse-DNS procedure, rate policy, or current canonical contact. The registry header alone is not authentication.
Compare observed behavior with a search-indexing hypothesis without turning it into attribution. Requests for public HTML, canonical metadata, feeds, sitemaps, and ordinary assets may be consistent with indexing. Private endpoint access, high concurrency, repeated retries, unexpected downloads, or traffic that ignores your policy may indicate spoofing, abuse, a fork, or a different client. These observations establish impact, not operator identity or downstream use.
Evaluate /robots.txt independently if you choose to publish a rule. Confirm that it is served from the intended host, returns a successful status and text content type, and contains the exact group you intend to apply. Test the complete header and any global group separately. Page-level directives can express discovery preferences:
<meta name="robots" content="noindex, nofollow">
X-Robots-Tag: noindex, nofollow
These signals do not authenticate a historical Istella client or secure private routes. Enforce sensitive boundaries in the application and at the origin.
WAF and Nginx remediation examples
If logs show a repeatable unwanted token, start in report-only mode and preserve representative requests. Adapt the expression to your WAF provider; it identifies a historical header pattern only:
{
"description": "Review historical IstellaBot traffic",
"expression": "lower(http.user_agent) contains \"istellabot\"",
"action": "log"
}
After reviewing false positives and confirming the business decision, scope enforcement to sensitive routes:
map $http_user_agent $review_istellabot {
default 0;
~*IstellaBot/1\.23\.15 1;
}
server {
location ~ ^/(admin|account|private|internal|api)/ {
if ($review_istellabot) { return 403; }
try_files $uri $uri/ =404;
}
}
Do not invent an IP allowlist, reverse-DNS suffix, or current operator exception. A User-Agent is easy to spoof and a version-specific historical match can catch an authorized compatibility client. Test public HTML, feeds, sitemaps, media, uploads, account flows, APIs, and approved integrations separately. Pair edge controls with authentication, rate limits, signed assets, caching, and anomaly detection.
Review checklist
Search logs for the complete registered header and preserve representative source IPs, ASNs, reverse-DNS results, paths, methods, response sizes, statuses, timing, and rate. Check whether the source can be independently verified and whether observed behavior resembles public search indexing. Do not classify a request as authentic from the token or community registry entry.
Re-check Istella, Tiscali, and the registry for a current operator source when an authorized accessible page is available. Keep this profile at legacy-label until a credible source publishes a current policy, canonical User-Agent, source-verification method, or robots guidance. Treat the Tiscali block and Google captcha as review limitations, not proof of inactivity.
Decide whether your objective is to preserve search visibility, limit extraction, protect private material, or reduce crawl load. Publish exact robots rules and enforce private routes with application and WAF controls. Do not claim successful blocking from configuration alone; verify subsequent access logs and response behavior.
References
- Tiscali homepage — registry-linked source; headless-browser access was blocked by policy restrictions on 2026-08-25.
- Google search attempt for IstellaBot — headless search was captcha-blocked on 2026-08-25; no result was used as evidence.
- Crawler User Agents community registry — community source for the historical User-Agent; it does not authenticate Istella infrastructure.
- Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
- RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.
Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.