Kangaroo Bot: Robots.txt & Crawl Policy Reference
Technical reference for the Kangaroo Bot label associated with the Australian Kangaroo LLM project. Learn how to verify historical or unidentified training traffic safely.
AI Summary: Kangaroo Bot is a registry label associated with the Australian Kangaroo LLM project and historical reports of AI-training data collection. The linked official crawler page is currently unavailable and resolves to a parked domain, so present operation, source ranges, the exact User-Agent, and robots compliance cannot be verified. Treat the token as an observation and enforce unwanted traffic with layered controls.
Role and policy boundary
Historical project reports describe Kangaroo LLM as an Australian AI initiative that planned to collect Australian English web content for model development. The registry consequently classifies Kangaroo Bot as an AI-training crawler. That context is useful for policy review, but it should not be presented as proof that a currently observed request is operated by Kangaroo LLM or that the project remains active.
The linked crawler page, kangaroollm.com.au/kangaroo-bot/, was not available during this review. Browser navigation redirected to a blank landing page, and direct extraction reported that the domain was parked by GoDaddy. No current first-party page could confirm a fixed User-Agent, source IP range, rate limit, contact address, or robots behavior.
If your logs confirm the exact token and you want to communicate a training-related restriction, publish:
User-agent: Kangaroo Bot
Disallow: /
This rule expresses your preference but does not prove that an unknown or legacy crawler will comply. Do not block unrelated browser traffic merely because the request resembles a historical Kangaroo Bot string. Protect private content with authentication and authorization.
Layered verification
Start with your own access logs. Search for Kangaroo Bot and preserve the complete User-Agent, source IP, ASN, reverse DNS, path, method, status, response size, timestamp, and rate. The registry-listed header is self-declared and may be stale or spoofed.
Examine whether requests resemble dataset collection: broad traversal across public pages, repeated retrieval of text and media, high concurrency, retries, or attempts to access image sitemaps and API endpoints. Behavioral evidence can establish resource impact, but it cannot establish the operator's model-training purpose.
Evaluate /robots.txt independently. Confirm the canonical host, response status, content type, exact group, and path match. Do not infer compliance from the absence of a block. For page-level indexing preferences, inspect:
<meta name="robots" content="noindex, nofollow">
X-Robots-Tag: noindex, nofollow
These directives communicate discoverability preferences and do not provide confidentiality. If a page must not be downloaded, require authentication or use protected object storage and signed URLs.
WAF and Nginx remediation examples
When logs show unwanted requests carrying the observed token, a narrow WAF rule can block that claim:
{
"description": "Block observed Kangaroo Bot token",
"expression": "lower(http.user_agent) contains \"kangaroo bot\"",
"action": "block"
}
For Nginx, scope enforcement to high-risk or high-cost routes first:
map $http_user_agent $block_kangaroo {
default 0;
~*Kangaroo[ -]Bot 1;
}
server {
location ~ ^/(api|uploads|originals|private|internal)/ {
if ($block_kangaroo) { return 403; }
try_files $uri $uri/ =404;
}
}
A User-Agent rule catches only the declared string, is easy to evade, and could produce false positives. Because no current source ranges were verified, do not add an IP allowlist based only on third-party reports. Pair edge rules with rate limits, authentication, signed URLs, and monitoring.
Review checklist
Search logs for the exact token and determine whether the traffic is present, active, high-volume, or limited to historical records. Check the official Kangaroo LLM domain before treating registry information as current; at this review, the linked crawler page was unavailable and parked. Preserve the date and evidence of that limitation.
Decide whether your objective is to prevent AI-training collection, reduce bandwidth, protect private content, or block a stale bot label. Publish a targeted robots group when it helps communicate the decision, then enforce high-risk routes with WAF, Nginx, application authorization, and rate limiting. Test public HTML, media, uploads, APIs, and authenticated pages separately. Revisit the profile if a current operator policy becomes available.
References
- Kangaroo Bot page — registry-linked operator URL; it was unavailable/parked during the 2026-08-24 review.
- Kangaroo LLM crawl report — third-party report describing the project's historical web-crawl announcement.
- Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.