Bot directory / ai-assistant

linkReader: Robots.txt & Crawl Policy Reference

Technical reference for the linkReader User-Agent associated with Sider's browser AI features. Learn how to distinguish user-directed fetching from a verified crawler contract.

AI Summary: linkReader/1.0 (+https://gochitchat.ai) is a browser-like token associated with Sider's AI tools for summarizing and working with web pages. Sider's official site confirms page reading and web automation features, but it does not publish a dedicated linkReader crawler contract, source range, or robots policy. Treat the token as a product-associated observation, not verified authentication, and use layered controls for unwanted access.

Role and policy boundary

Sider presents itself as an AI agent for the browser. Its Chat feature can summarize, explain, translate, and explore web pages, PDFs, and YouTube content; its Claw feature can research, compare, extract, and organize information across websites and operate on real sites. Those capabilities provide a plausible context for a link-reading request made on behalf of a user.

That product context does not prove that every request carrying linkReader/1.0 originates from Sider or that Sider always uses this header. The registry links the token to GoChitChat and Sider, but Sider's public product page does not publish a current source range, fixed crawl schedule, verification method, or robots compliance statement. A browser-like User-Agent can be changed by the client and copied by unrelated automation.

This profile therefore treats linkReader as a user-directed assistant token with partial documentation. Do not label it a search-index crawler or training crawler without a first-party statement. Blocking it affects requests declaring that token; it does not remove your pages from general search or block Sider workflows using another client identity.

If your logs confirm the token and you want to communicate a restriction, publish:

configuration / code
User-agent: linkReader
Disallow: /

For selective access to public reference material:

configuration / code
User-agent: linkReader
Allow: /docs/
Allow: /public-reference/
Disallow: /private/
Disallow: /account/
Disallow: /checkout/

Robots.txt is advisory. Protect private and transactional routes with authentication and authorization.

Layered verification

Start with access logs and preserve the complete User-Agent, source IP, ASN, reverse DNS, path, method, status, response size, redirect chain, timestamp, and request rate. A match for linkReader is a classification clue, not proof of Sider origin.

Compare traffic with Sider's current product behavior. User-directed page reading may produce isolated requests for a page a user selected, while automated research or Claw tasks may follow several related links. These patterns help with risk assessment but do not establish a formal service contract. No authoritative linkReader IP range or reverse-DNS process was found in the public documentation reviewed.

Evaluate /robots.txt independently. Confirm the canonical host, status, content type, exact group, and path match. Page-level metadata can express discoverability preferences:

configuration / code
<meta name="robots" content="noindex, nofollow">
configuration / code
X-Robots-Tag: noindex, nofollow

A user-triggered reader may still need to retrieve a noindex page to answer the user's request, and metadata does not provide confidentiality. Use route authorization, signed URLs, and rate limits for sensitive content.

WAF and Nginx remediation examples

If logs confirm unwanted requests carrying the token, apply a narrow WAF rule:

configuration / code
{
  "description": "Block observed linkReader token",
  "expression": "lower(http.user_agent) contains \"linkreader\"",
  "action": "block"
}

For Nginx, restrict enforcement to high-risk routes while investigating:

configuration / code
map $http_user_agent $block_linkreader {
    default 0;
    ~*linkReader 1;
}

server {
    location ~ ^/(private|account|checkout|internal|api)/ {
        if ($block_linkreader) { return 403; }
        try_files $uri $uri/ =404;
    }
}

A header rule can be spoofed, rotated, or absent from another Sider client path. Test against ordinary browsers, social previews, search crawlers, and approved Sider workflows. Combine edge matching with authentication, rate limits, and behavior-based bot controls rather than treating the header as identity proof.

Review checklist

Search logs for linkReader/1.0 and record source, volume, paths, response sizes, and timing. Re-check Sider's current product and help pages, and note that the reviewed public page documents browser reading and automation but not a separate linkReader crawler contract.

Decide whether your goal is to support user-directed research, preserve public visibility, protect private pages, or reduce extraction load. Publish a targeted robots group if it communicates the decision, then enforce high-risk routes with WAF, Nginx, authentication, and rate limiting. Test public docs, private pages, APIs, and media independently. Revisit the profile if Sider publishes a dedicated User-Agent, source-verification, or robots policy.

References

  1. Sider AI — official product page describing browser-based reading, research, extraction, and automation features.
  2. GoChitChat AI — registry-linked product domain; captcha-protected during this review, so no crawler policy was confirmed there.
  3. Google Robots.txt Introduction — general explanation of crawler directives and their limitations.

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.