← Bot Directory/TSM-turingos
Bot directory / ai-training

TSM-turingos: Robots.txt & Crawl Policy Reference

Technical reference for the TSM-turingos User-Agent label. Learn how to investigate unidentified AI-related traffic when no current operator crawler contract is available.

AI Summary: TSM-turingos is a registry label with a browser-like User-Agent, while the linked Known Agents page classifies turingos as an uncategorized agent with no established purpose or verification method. No current first-party TuringOS crawler policy was found. Treat the header as a self-declared observation and use logs, robots communication, and active controls to manage unwanted traffic.

Role and policy boundary

The registry describes TSM-turingos as a TuringOS AI-powered content-analysis and monitoring crawler. The available third-party reference does not support that level of certainty. Known Agents calls turingos uncategorized, says it has no established behavior profile or identified purpose, and advises owners to infer risk from their own request patterns.

The observed header contains TSM-turingos-1253296984, but a browser-compatible User-Agent can be copied or changed. No first-party TuringOS source reviewed for this profile publishes the token, source ranges, crawl schedule, purpose, or robots contract. Do not claim that the traffic is collecting training data, indexing content, or complying with robots.txt from the registry label alone.

If your logs confirm the exact observed token and you want to communicate a restriction, use the stable part carefully:

configuration / code
User-agent: TSM-turingos
Disallow: /

For selective access:

configuration / code
User-agent: TSM-turingos
Allow: /public-reference/
Allow: /docs/
Disallow: /private/
Disallow: /internal/
Disallow: /api/

Robots.txt is advisory. It cannot protect publicly reachable private content; use authentication and authorization for sensitive routes.

Layered verification

Start with raw logs and preserve the complete User-Agent, source IP, ASN, reverse DNS, method, path, status, response size, redirects, timestamp, and request rate. Because the directory says there is no reliable verification method, treat a header match as a clue, not proof of operator identity.

Analyze behavior without assigning purpose prematurely. Steady crawling may resemble an indexer; fast, deep sweeps may resemble scraping; repeated requests to login, admin, or API paths may indicate abuse. These patterns establish operational risk but cannot prove a TuringOS operator or AI-training use.

Do not treat the Known Agents country, statistics, or expected robots behavior as authoritative source-network or compliance evidence. Evaluate your own observations and seek a first-party response before allowlisting. Inspect /robots.txt independently for status, content type, exact group, path match, and host scope.

Page-level metadata may express a discoverability preference:

configuration / code
<meta name="robots" content="noindex, nofollow">
configuration / code
X-Robots-Tag: noindex, nofollow

These directives do not provide confidentiality and may be ignored by an undocumented client. Protect private content with application authorization, signed URLs, and a protected origin.

WAF and Nginx remediation examples

Once logs confirm an unwanted token, a narrow WAF rule can block the declared identity:

configuration / code
{
  "description": "Block observed TSM-turingos token",
  "expression": "lower(http.user_agent) contains \"tsm-turingos\"",
  "action": "block"
}

For Nginx, scope enforcement to high-risk routes while investigating public pages:

configuration / code
map $http_user_agent $block_tsm_turingos {
    default 0;
    ~*TSM-turingos 1;
}

server {
    location ~ ^/(private|internal|account|uploads|api)/ {
        if ($block_tsm_turingos) { return 403; }
        try_files $uri $uri/ =404;
    }
}

A User-Agent rule is easy to spoof or evade and can block an approved client using the same substring. Do not create an IP allowlist without a verified operator-published range. Test browsers, social previews, feed readers, search crawlers, and monitors. Pair edge matching with authentication, rate limits, signed assets, and anomaly detection.

Review checklist

Search logs for TSM-turingos and record the full header, source network, paths, response sizes, statuses, and timing. Keep the registry description separate from verified evidence. Re-check the Known Agents page and look for a first-party TuringOS source; during this review no operator policy or reliable verification method was available.

Decide whether your objective is to prevent possible content collection, protect bandwidth, secure private routes, or investigate unexplained requests. Publish a targeted robots group for the exact observed token, enforce sensitive routes with WAF and application controls, and test docs, media, feeds, sitemaps, uploads, and APIs separately. Revisit the profile if the operator publishes a canonical User-Agent, source verification, purpose statement, or robots policy.

References

  1. Known Agents turingos entry — third-party classification stating that the agent is uncategorized and lacks a reliable verification method.
  2. Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
  3. RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.