Bot directory / search-engine

Teoma: Robots.txt & Crawl Policy Reference

Technical reference for the historical Teoma (Ask Jeeves) registry label, noting the official shutdown of Ask.com in May 2026.

AI Summary: The registry describes Teoma as the Ask Jeeves search engine crawler and records Mozilla/2.0 (compatible; Ask Jeeves/Teoma; +http://sp.ask.com/docs/about/tech_crawling.html). However, IAC officially discontinued its search business and closed Ask.com on May 1, 2026. The crawler documentation URLs no longer resolve. Treat the token as an obsolete historical label rather than an active search crawler.

Role and policy boundary

The registry describes this entry as the Ask Jeeves Teoma search engine web crawler and records this historical User-Agent:

configuration / code
Mozilla/2.0 (compatible; Ask Jeeves/Teoma; +http://sp.ask.com/docs/about/tech_crawling.html)

During headless-browser review on August 25, 2026, the registry-linked documentation URLs (about.ask.com and sp.ask.com) failed to resolve. The root domain www.ask.com displayed an official farewell message stating that IAC has discontinued its search business and that Ask.com officially closed on May 1, 2026.

This profile therefore uses legacy-label and does not present the registry description as a current operator policy. A matching request is no longer an authentic Ask.com crawler; it could be a legacy deployment that was never decommissioned, a third-party scraper, a test client, or a spoofed header. Do not infer that an active Ask Jeeves search crawler still exists, that it obeys robots.txt, or that the token establishes permission for AI input, model training, or data retention.

Because the service is discontinued, do not publish an assumed Teoma robots group as if it came from the operator. If logs establish a current, attributable token and you decide to exclude it, use the exact observed value in a deliberate site-owner rule:

configuration / code
User-agent: CONFIRMED-OBSERVED-TEOMA
Disallow: /

For selective access after confirmation:

configuration / code
User-agent: CONFIRMED-OBSERVED-TEOMA
Allow: /public/
Allow: /docs/
Disallow: /admin/
Disallow: /account/
Disallow: /private/
Disallow: /licensed/
Disallow: /api/

These are explanatory site-owner controls, not recovered Ask.com instructions. Robots.txt is advisory and cannot secure private or licensed content; use authentication, authorization, signed URLs, data minimization, and origin controls.

Layered verification

Start with raw access logs and preserve the complete User-Agent, source IP, ASN, reverse-DNS result, HTTP method, requested path, response status, response size, redirect chain, timestamp, concurrency, and request rate. Because Ask.com has shut down, no current source publishes an IP range, reverse-DNS procedure, Crawl-delay, rate guidance, removal address, or verification workflow.

Do not treat the ask.com URL in a header as proof of operator ownership. A client can copy any URL, and the referenced subdomains no longer resolve. Check source IP and DNS evidence independently. An IP that historically belonged to IAC or Ask.com is not proof of an authorized current crawler.

Compare observed behavior with a search-indexing hypothesis without turning it into attribution. Public HTML, metadata, feeds, sitemaps, and ordinary assets may be requested by many tools. Private endpoints, licensed content, account routes, APIs, high concurrency, repeated retries, or unexpected bulk downloads establish impact and load risk, not Teoma attribution or downstream use.

Evaluate /robots.txt only after the actual observed token is known. Confirm that it is served by the intended host, returns a successful text response, and contains the exact group you intend to publish. An absent operator source and a discontinued service are not evidence of robots compliance. If no exact group exists, a global rule may affect unrelated clients and should be adopted only as an explicit site-wide decision.

Page-level directives can express indexing preferences:

configuration / code
<meta name="robots" content="noindex, nofollow">
configuration / code
X-Robots-Tag: noindex, nofollow

These signals do not authenticate an undocumented crawler or secure private paths. Enforce sensitive boundaries in the application and at the origin. If your policy distinguishes search indexing, AI input, reference use, and model training, document each purpose separately rather than inferring permission from a discontinued registry label.

WAF and Nginx remediation examples

When the source is unverified, use report-only logging and a narrow observation. Avoid broad ask, teoma, or browser rules that can affect unrelated clients:

configuration / code
{
  "description": "Observe unverified Teoma candidates",
  "expression": "lower(http.user_agent) contains \"teoma\"",
  "action": "log"
}

After independent verification and a policy decision, scope enforcement to sensitive routes and preserve evidence for the rule:

configuration / code
map $http_user_agent $block_confirmed_teoma_private {
    default 0;
    # Add only a complete, independently verified token here.
    # ~*Teoma 1;
}

server {
    location ~ ^/(admin|account|private|licensed|internal|api)/ {
        if ($block_confirmed_teoma_private) { return 403; }
        try_files $uri $uri/ =404;
    }
}

A User-Agent match is easy to spoof, and the registry URL points to a discontinued service. Do not invent an IP allowlist, reverse-DNS suffix, rate, training policy, or permanent trust exception. Use per-client rate limits, concurrency ceilings, timeouts, response-size controls, caching, and anomaly detection at the edge or origin. Start in report-only mode, review false positives, and restrict only after evidence supports the action.

Test public pages, feeds, sitemaps, structured data, licensed assets, account routes, APIs, 429 behavior, response-size limits, and approved integrations separately. Pair WAF controls with authentication and application authorization instead of using robots.txt as an access-control mechanism.

Review checklist

Search logs for the exact Teoma substring and preserve the complete header, source IP, ASN, PTR result, forward lookup, path, method, response size, status, timing, rate, concurrency, and redirects. Treat the URL in the header as an unverified string, not a support channel.

Re-check the Ask.com domain, robots.txt, and the crawler registry. In this review, Ask.com explicitly confirmed its shutdown on May 1, 2026. Keep this profile at legacy-label permanently unless the brand is verifiably resurrected by a new operator.

Decide whether your objective is to preserve public discovery, reduce crawl load, limit extraction, protect licensed material, or prevent private access. Publish exact robots rules only after the token is known; enforce private routes with authentication and origin controls.

Review search indexing, AI input, reference use, and model-training decisions separately. Neither the registry description nor a header containing an obsolete URL establishes downstream permission or use. Do not claim successful blocking or verification from configuration alone; validate later logs.

Record the date and reason for the legacy-label classification so future evidence can be compared without silently upgrading an unsupported historical label. If a new deployment identifies itself differently, create a separate evidence trail.

References

  1. Crawler User Agents community registry — registry context for the Teoma label and User-Agent; it does not authenticate current infrastructure.
  2. Ask.com homepage — root domain; headless-browser review on 2026-08-25 confirmed the service was discontinued on May 1, 2026.
  3. Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
  4. RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.