← Bot Directory/imageSpider
Bot directory / ai-training

imageSpider: Robots.txt & Crawl Policy Reference

Technical reference for the imageSpider User-Agent token associated with ByteDance in third-party registries. Learn how to verify image collection traffic without assuming a training contract.

AI Summary: imageSpider is a browser-compatible User-Agent token associated with ByteDance by third-party bot directories, but no current first-party ByteDance crawler policy was found. The token, image-collection purpose, source ranges, and robots compliance are therefore not fully verified. Treat it as an observed identifier, inspect image-request behavior in your logs, and use narrow robots and WAF controls only when the token is actually present.

Role and policy boundary

Third-party crawler references describe imageSpider as a ByteDance-associated image crawler. Reports also discuss unusual ByteDance request headers and image collection, but those sources do not establish a current, operator-controlled contract. ByteDance's public corporate and model pages reviewed for this profile do not publish a dedicated imageSpider page that confirms its User-Agent, IP ranges, crawl rate, data use, or robots behavior.

This means the apparent vendor and training-related purpose must remain qualified. A request carrying imageSpider may be genuine, stale, or spoofed. Do not claim that every imageSpider request contributes to model training, and do not assume that a robots.txt allow result means the operator will comply.

If your own logs confirm the token and your policy is to communicate a no-access preference, publish:

configuration / code
User-agent: imageSpider
Disallow: /

If you want to expose selected public media while protecting originals and user uploads:

configuration / code
User-agent: imageSpider
Allow: /public-images/
Allow: /thumbnails/
Disallow: /originals/
Disallow: /uploads/
Disallow: /api/

Robots rules express a preference; they do not protect private media. Use object authorization, signed URLs, and origin controls for restricted assets.

Layered verification

Start with the exact User-Agent from the registry and compare it with observed logs. Preserve source IP, ASN, reverse DNS, request path, method, status, response size, referer, timestamp, and rate. The browser-like prefix may be shared with unrelated automation, and a legitimate collector may use another header.

Pay special attention to media behavior: requests for full-resolution files, image variants, CSS sprite sheets, upload directories, image sitemaps, or API endpoints. Repeated high-rate image fetches can establish operational risk even when the crawler's purpose remains uncertain.

No authoritative imageSpider source ranges or reverse-DNS verification method were found. Treat network information as corroborating telemetry, not proof. Evaluate /robots.txt independently by checking the canonical host, status, content type, exact group, and path match. For indexing preferences, inspect:

configuration / code
<meta name="robots" content="noindex, nofollow">
configuration / code
X-Robots-Tag: noindex, nofollow

Metadata can signal discoverability preferences but cannot substitute for media access controls or stop an unknown crawler from downloading a response.

WAF and Nginx remediation examples

When logs show unwanted traffic that declares imageSpider, a narrow WAF expression can block that claim:

configuration / code
{
  "description": "Block observed imageSpider token",
  "expression": "lower(http.user_agent) contains \"imagespider\"",
  "action": "block"
}

For Nginx, start with high-cost media paths rather than blocking all browser traffic:

configuration / code
map $http_user_agent $block_imagespider {
    default 0;
    ~*imageSpider 1;
}

server {
    location ~ ^/(images|media|uploads|originals|api)/ {
        if ($block_imagespider) { return 403; }
        try_files $uri $uri/ =404;
    }
}

Test against browsers, image optimization services, social previews, and internal monitors. A header rule is easy to evade and can produce false positives. Pair it with rate limiting, hotlink controls, signed URLs, and authenticated access for sensitive media. Do not create an IP allowlist until ByteDance publishes verifiable current ranges.

Review checklist

Search access logs for imageSpider and confirm whether it requests images, HTML pages, or protected routes. Separate evidence from attribution: record the header and behavior, but do not label the traffic as training activity without a first-party statement. Review the third-party directory and ByteDance's current public documentation for changes.

Decide whether the goal is to prevent training-related extraction, protect bandwidth, stop image harvesting, or restrict sensitive assets. Publish the specific robots group if communication is useful, then enforce high-risk decisions with WAF and application controls. Test public images, originals, uploads, image sitemaps, and API endpoints independently. Revisit the profile when ByteDance publishes a dedicated imageSpider policy or verification method.

References

  1. KnownAgents imageSpider entry — third-party identification reference; it is not an operator-controlled policy.
  2. ByteDance — linked company domain; no dedicated imageSpider crawler contract was identified during this review.
  3. Google Robots.txt Introduction — general explanation of crawler directives and their limitations.

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.