← Bot Directory/Spawning-AI
Bot directory / ai-training

Spawning-AI: Robots.txt & Crawl Policy Reference

Technical reference for the Spawning-AI registry label. Learn how to investigate possible AI-training traffic when no current operator crawler contract is available.

AI Summary: Spawning-AI is an inventory label associated with possible AI-training or content-collection traffic, but the registry-linked reference currently returns a DataDome 404 page and no current operator crawler contract could be verified. Its User-Agent, source network, purpose, and robots behavior are unknown. Treat any observed request as evidence to investigate, not proof of attribution.

Role and policy boundary

The registry categorizes Spawning-AI as an AI-training data collector, but no current operator-controlled source was found that establishes this identity or purpose. The linked DataDome reference is unavailable and does not act as a policy for a Spawning-AI operator. A label in a crawler inventory cannot prove that a request is collecting training data, operating a search index, or honoring a site's robots file.

No canonical User-Agent, source IP list, rate policy, contact address, or opt-out contract is publicly established in the sources reviewed. If your logs show a related string, preserve the exact header and do not normalize it automatically to Spawning-AI. A client can claim any header and may rotate it between jobs.

If you confirm an exact token and want to communicate a restriction, replace the placeholder group with the string observed in logs:

configuration / code
User-agent: Spawning-AI
Disallow: /

For a selective policy:

configuration / code
User-agent: Spawning-AI
Allow: /public-reference/
Allow: /docs/
Disallow: /private/
Disallow: /uploads/
Disallow: /api/

Robots.txt is advisory. It cannot protect publicly reachable private or licensed content; use authentication, authorization, and contractual controls where appropriate.

Layered verification

Start with raw access logs. Record the complete User-Agent, source IP, ASN, reverse DNS, method, path, status, response size, redirects, timestamp, and request rate. Because the identity is undocumented, do not rely on a partial substring or a third-party list as authentication.

Assess behavior without over-interpreting it. Broad retrieval of HTML, media, sitemaps, and feeds may resemble dataset collection; repeated downloads of original assets may indicate bandwidth or licensing risk; requests for private APIs may indicate an access-control problem. These are operational signals, not proof of an operator's purpose.

Evaluate /robots.txt independently. Confirm the canonical host, response status, content type, exact user-agent group, and path match. Page-level metadata may express a discoverability preference:

configuration / code
<meta name="robots" content="noindex, nofollow">
configuration / code
X-Robots-Tag: noindex, nofollow

These directives do not establish a training opt-out for an undocumented client and do not secure private assets. Use authenticated delivery, signed URLs, and origin controls for sensitive material.

WAF and Nginx remediation examples

Once logs confirm an exact unwanted token, a narrow WAF rule can block the self-declared value:

configuration / code
{
  "description": "Block observed Spawning-AI-related token",
  "expression": "lower(http.user_agent) contains \"spawning-ai\"",
  "action": "block"
}

For Nginx, scope enforcement to private and high-cost paths while you investigate public content:

configuration / code
map $http_user_agent $block_spawning_ai {
    default 0;
    ~*spawning-ai 1;
}

server {
    location ~ ^/(private|uploads|originals|internal|api)/ {
        if ($block_spawning_ai) { return 403; }
        try_files $uri $uri/ =404;
    }
}

A User-Agent rule is easy to spoof or evade and may block an approved research process. Do not build an IP allowlist without a verified operator-published range. Test browsers, image optimizers, feed readers, social previews, and approved integrations. Pair edge matching with authentication, rate limits, signed assets, and monitoring.

Review checklist

Search logs for every exact header that may be associated with the label and preserve representative requests. Record request volume, paths, response sizes, source networks, status, and timing. Re-check the registry-linked reference and look for a current operator page; during this review the DataDome URL returned a 404 and no policy was available.

Decide whether your objective is to prevent possible model-data collection, protect licensed media, reduce bandwidth, or secure private routes. Publish a targeted robots group only for the exact observed token, enforce sensitive content with application controls, and test public docs, feeds, sitemaps, media, uploads, and APIs separately. Revisit the profile if the operator publishes a canonical User-Agent, source verification, purpose statement, or opt-out process.

References

  1. Registry-linked Spawning bot reference — returned a DataDome 404 during the 2026-08-24 headless-browser review; it is not treated as operator evidence.
  2. DataDome Web & LLM scraping resources — general anti-scraping context, not a Spawning-AI policy.
  3. Google Robots.txt Introduction — general explanation of crawler directives and their limitations.

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.