← Bot Directory/ChatGLM-Spider
Bot directory / ai-training

ChatGLM-Spider: Robots.txt & Crawl Policy Reference

Technical reference for the registry-listed ChatGLM-Spider User-Agent. Learn what Zhipu AI publicly documents, how to verify the crawler, and how to protect site content.

AI Summary: ChatGLM-Spider/1.0 is a registry-listed crawler associated with the ChatGLM/Zhipu AI ecosystem. Zhipu AI’s public developer documentation describes GLM models, web search, knowledge-base retrieval, and agent development, but the official sources reviewed here do not publish a definitive crawler contract or robots.txt compliance statement for this exact token. Treat the User-Agent as an observed identifier, verify it in logs, and do not use it as authentication.

Role and policy boundary

The name ChatGLM-Spider suggests a crawler related to the ChatGLM ecosystem, and the project registry classifies it as an AI-training data collection crawler. That classification is useful for triage, but it should not be promoted to a vendor-confirmed fact without a primary crawler document. Zhipu AI’s official platform documentation describes a broad developer platform: GLM model invocation, web search, knowledge-base retrieval, and agent capabilities. Those product capabilities do not by themselves prove how a public web request is collected, stored, indexed, or used.

For a webmaster, the correct decision is therefore evidence-led. If the request is observed and the organization does not want unidentified AI-oriented collection, a dedicated robots group can express a clear opt-out. If public pages are intentionally offered to downstream search or knowledge workflows, allow only the paths that are meant to be public. Keep private documentation, user records, staging environments, and transactional endpoints behind authentication regardless of what a robots file says.

A robots rule is a published crawl preference; it does not authenticate a caller or enforce confidentiality. The User-Agent is self-declared and can be copied by unrelated software. Do not allow a request to cross an application authorization boundary solely because it carries the ChatGLM-Spider token.

To allow only a public documentation area after verifying the request in logs, use:

configuration / code
User-agent: ChatGLM-Spider
Allow: /docs/
Disallow: /staging/
Disallow: /internal/
Disallow: /account/

To request that the identified token not crawl the site, use:

configuration / code
User-agent: ChatGLM-Spider
Disallow: /

Avoid using a wildcard when the decision is specific to this crawler. A wildcard can unintentionally change the behavior of search engines, user-triggered assistants, and other training crawlers.

Layered verification

Begin with the complete request evidence: User-Agent, source address, request path, HTTP method, status, redirect chain, and request rate. Compare the observed token with the current registry value and with any new Zhipu AI documentation. The exact ChatGLM-Spider/1.0 token is not independently confirmed by the official product and developer pages reviewed for this profile, so the audit should label it as a registry observation rather than a verified vendor identity.

Then evaluate the top-level /robots.txt response and the path-specific directive. Confirm that the file is served from the canonical host, is publicly reachable, and is not replaced by a login page, challenge, or redirect to a different policy. Inspect the returned HTML and headers separately. A robots.txt allow result means only that the published crawl preference permits the path; it does not demonstrate that the caller is operated by Zhipu AI or that the content will be used in a particular way.

The Policy Engine should keep these layers separate. A User-Agent match can explain why a rule was selected, while the lack of a primary vendor crawler specification should remain visible as an evidence limitation. If the request is sensitive, require origin authentication and apply source verification or signed-request controls where your infrastructure supports them.

Page-level directives communicate indexing preferences, but do not replace access control:

configuration / code
<meta name="robots" content="noindex, nofollow">
configuration / code
X-Robots-Tag: noindex, nofollow

Use these signals for discoverability and indexing policy. Use authentication, authorization, WAF controls, and rate limiting to protect private data and server resources.

WAF and Nginx remediation examples

If your policy is to block requests that self-identify as ChatGLM-Spider, match the stable token and record the decision as a network-layer policy. This is not an identity check and may also catch spoofed requests:

configuration / code
{
  "description": "Block declared ChatGLM-Spider token",
  "expression": "lower(http.user_agent) contains \"chatglm-spider\"",
  "action": "block"
}

A narrow Nginx example for protected paths is:

configuration / code
map $http_user_agent $deny_chatglm_spider {
    default 0;
    ~*chatglm-spider 1;
}

server {
    location ~ ^/(internal|account|customer-data)/ {
        if ($deny_chatglm_spider) { return 403; }
        try_files $uri $uri/ =404;
    }
}

Test this in staging and do not rely on a User-Agent match as the only protection for confidential content. If later vendor documentation confirms a different token or a signed identity mechanism, update the policy and the registry evidence rather than silently widening the block.

Review checklist

Request /robots.txt from the canonical host and verify its status, content type, final URL, and dedicated ChatGLM-Spider group. Test one allowed documentation URL, one disallowed internal URL, and one protected account URL. Capture the complete User-Agent, source address, HTTP status, redirect chain, response headers, and request rate.

Finally, compare the observed behavior with current Zhipu AI documentation and keep the evidence classification current. Do not claim that the crawler respects robots.txt, trains a specific model, or represents Zhipu AI traffic unless a primary source or verifiable network evidence supports that claim. Re-test after CDN, WAF, origin, or content-policy changes.

References

  1. Zhipu AI ChatGLM official site — product context for ChatGLM.
  2. Zhipu AI developer platform introduction — official documentation for GLM platform capabilities, web search, knowledge retrieval, and agents.

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.