LinkupBot: Robots.txt & Crawl Policy Reference
Technical reference for LinkupBot, Linkup's web-indexing crawler for AI search and grounding. Learn its User-Agent, CIDR verification, and robots controls.
AI Summary: LinkupBot is Linkup's documented web crawler for building and maintaining the Linkup search index, which powers search and web grounding for AI applications. Linkup publishes the
LinkupBottoken, a complete User-Agent example, a changing CIDR list, robots.txt examples, and a dual-signal verification method. Use the exact robots group for search visibility decisions and verify both the source IP and declared token before allowlisting.
Role and policy boundary
Linkup describes LinkupBot as the crawler behind its web search index. It discovers and retrieves publicly accessible webpages so Linkup can provide search results and web-grounded context to AI applications. This is a recurring indexing service rather than a one-off browser fetch, so allowing it may make public documentation available through Linkup-powered applications.
The official page publishes the token LinkupBot and the example LinkupBot/1.0 (LinkupBot for web indexing; https://www.linkup.so/bot; bot@linkup.so). Linkup also states that the crawler respects applicable robots.txt rules. Blocking LinkupBot changes Linkup indexing and grounding availability; it does not block ordinary search engines or other AI crawlers.
For a site that wants to allow public documentation but exclude private areas:
User-agent: LinkupBot
Allow: /docs/
Allow: /public-reference/
Disallow: /private/
Disallow: /drafts/
Disallow: /internal-document.pdf
To block the crawler site-wide:
User-agent: LinkupBot
Disallow: /
To allow LinkupBot while blocking other crawlers, Linkup documents an explicit group order:
User-agent: *
Disallow: /
User-agent: LinkupBot
Allow: /
A WAF allowlist does not override a disallow rule. Protect confidential content with authentication and server-side authorization.
Layered verification
Linkup explicitly warns that User-Agent headers can be spoofed. Its recommended verification requires two signals: the source IP must fall within the current CIDR ranges published at https://www.linkup.so/linkupbot-ips.txt, and the User-Agent must contain the LinkupBot token. Do not allowlist a request based on the header alone.
Record the full User-Agent, source IP, reverse DNS, ASN, request path, method, status, response size, redirect chain, timestamp, and request rate. Refresh the published CIDR list because Linkup states that it may change. Avoid matching the complete versioned User-Agent string; the token is more stable than the version and explanatory suffix.
Evaluate /robots.txt independently for the same host, protocol, and port. Linkup's documentation states that separate subdomains may need separate files. Confirm response status, content type, exact group, path matching, and whether a CDN or application is returning the intended policy.
For page-level discoverability, inspect:
<meta name="robots" content="noindex, nofollow">
X-Robots-Tag: noindex, nofollow
These directives are separate from Linkup's crawler policy and do not protect private routes. Use authentication and authorization for confidential data.
WAF and Nginx remediation examples
If you want to permit verified LinkupBot traffic, require both the published network range and the token. The CIDR values below are intentionally not hard-coded because Linkup says its list may change; retrieve the current list through your controlled configuration process.
For a blocking WAF rule, match the token only when you have decided to remove Linkup search visibility:
{
"description": "Block LinkupBot search crawler",
"expression": "lower(http.user_agent) contains \"linkupbot\"",
"action": "block"
}
For Nginx, a token match can enforce a simple block, while an allowlist should be maintained from Linkup's current CIDR file:
map $http_user_agent $is_linkupbot {
default 0;
~*LinkupBot 1;
}
server {
location / {
if ($is_linkupbot) { return 403; }
try_files $uri $uri/ =404;
}
}
Do not use this header-only example as proof of identity. For an allow policy, combine the token with a current source-IP allowlist, then still ensure /robots.txt permits the requested path. Test changes in staging and include the complete request details when contacting bot@linkup.so about suspected misbehavior.
Review checklist
Decide whether Linkup search and AI grounding are valuable for your public content. Publish a host-specific User-agent: LinkupBot group that allows, selectively allows, or blocks the intended paths. Remember that the WAF decision and robots decision are separate: both must permit access for a request to be useful.
For verification, compare the complete User-Agent with the current CIDR list at linkupbot-ips.txt. Check reverse DNS and forward confirmation where appropriate, record status and response size, and monitor request rate. Do not match a complete version string because it may change. Protect private pages, APIs, accounts, and documents with authentication regardless of crawler policy.
If traffic appears abusive, preserve the URL, timestamp and timezone, source IP, complete User-Agent, and relevant logs, then contact Linkup at bot@linkup.so. Re-test public docs, excluded paths, subdomains, and protected routes after CDN or WAF changes.
References
- LinkupBot — official crawler role, User-Agent, robots examples, IP verification, and contact guidance.
- LinkupBot IP ranges — official changing CIDR list for source verification.
- Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.