Bot directory / AI Training Crawler

Bytespider — ByteDance crawler

A technical reference for Bytespider and training-oriented crawler controls.

AI Summary: Bytespider is grouped with training-oriented crawlers in this seed catalog. Its presence in a user-agent rule should be reported as declared policy evidence, not a guaranteed enforcement result.

Role and policy boundary

Treat this User-Agent as a distinct policy identity. The correct decision depends on the bot purpose, the paths it requests, and the response controls applied by the origin, CDN, and application layer. A robots rule is a declaration of intent; it does not replace authentication, authorization, or rate limiting.

Bytespider is treated as a training-oriented crawler in this seed catalog, with an observed status rather than a claim of permanent vendor behavior. That label is useful for policy planning, but it should remain visible as a catalog assumption. Keep the catalog version and verification date beside the decision so later audits can distinguish a changed vendor from a changed site policy.

To block the crawler everywhere, use:

configuration / code
User-agent: Bytespider
Disallow: /

To allow public resources while excluding sensitive areas:

configuration / code
User-agent: Bytespider
Allow: /public/
Disallow: /

When the same file contains several user-agent blocks, the engine evaluates the matching block and wildcard behavior for the chosen path. A malformed line or empty directive is retained as incomplete evidence rather than silently treated as a valid allow.

Layered verification

Verify the same URL through each control plane instead of assuming that one green signal represents the whole request path. Compare the bot-specific robots group, the page-level metadata, and the response headers captured at the public edge.

Robots.txt is not the only signal. A page-level directive can restrict indexing or retrieval:

configuration / code
<meta name="robots" content="noindex, nofollow">
configuration / code
X-Robots-Tag: noindex

If robots.txt permits Bytespider but the response sends noindex, the result is a cross-layer conflict. If the page is unavailable or the response is too large to inspect, the diagnostic reports partial or unknown evidence. It should never call that situation a clean pass.

WAF and Nginx remediation examples

An advisory Cloudflare rule can identify Bytespider and log requests before enforcement:

configuration / code
{
  "description": "Review Bytespider training access",
  "expression": "lower(http.user_agent) contains \"bytespider\"",
  "action": "log",
  "note": "Confirm the policy and test cache behavior before blocking."
}

A Nginx map can classify and block after the owner has approved the change:

configuration / code
map $http_user_agent $bytespider_action {
    default allow;
    ~*"Bytespider" block;
}

Do not insert a generated snippet into a production WAF without syntax validation, staged rollout, and an emergency rollback path. The generated content is intentionally advisory and makes no guarantee about crawler compliance.

Review checklist

Use this checklist after every policy change and after a CDN or WAF migration. Record the request URL, User-Agent, HTTP status, final redirect, and the exact evidence used to reach the decision.

Check exact user-agent matching, public/private path scope, page metadata, response headers, CDN behavior, and whether a redirect lands on a different hostname. Record evidence and catalog version. If the vendor’s behavior is uncertain, say so in the report and use an operational control such as authentication for truly sensitive content.


Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.