← Bot Directory/kagi-fetcher
Bot directory / ai-assistant

kagi-fetcher: Robots.txt & Crawl Policy Reference

Technical reference for the registry-linked kagi-fetcher token associated with Kagi AI. Learn why it must be distinguished from Kagi's documented Kagibot crawler.

AI Summary: kagi-fetcher/1.0 is a registry-linked token associated with Kagi's AI features, but Kagi's official public crawler page documents a different bot: Kagibot. Kagi's AI documentation confirms assistant and page-summarization features but does not establish a separate kagi-fetcher User-Agent, source range, or robots contract. Treat this token as an unverified product-associated observation and do not apply Kagibot's policy to it.

Role and policy boundary

Kagi's official documentation describes Kagi Assistant, Quick Answer, Summarize Page, and question-answering features that use AI within search and page contexts. These features provide a plausible product context for a fetcher that retrieves a page on behalf of an end user, but the public documentation reviewed for this profile does not publish a separate kagi-fetcher crawler specification.

Kagi does publish a dedicated page for Kagibot, its search-engine crawler. That page gives Kagibot's exact browser-compatible User-Agent, four source IPs, and robots behavior. kagi-fetcher/1.0 is not the same token. Do not reuse Kagibot's IP allowlist, compliance statement, or search-indexing semantics for the registry-linked fetcher.

This distinction also affects policy decisions. A user-triggered assistant fetch can be different from a scheduled search crawl, and blocking one token does not remove a page from Kagi Search or prevent all Kagi products from requesting it. Because no first-party kagi-fetcher contract was found, do not claim that it respects robots, ignores robots, or uses a fixed source network.

If your logs confirm this exact token and you want to communicate a restriction, publish:

configuration / code
User-agent: kagi-fetcher
Disallow: /

For a selective policy:

configuration / code
User-agent: kagi-fetcher
Allow: /public-reference/
Allow: /docs/
Disallow: /account/
Disallow: /checkout/
Disallow: /private/

Treat robots.txt as a preference. Use authentication and authorization for confidential or transactional content.

Layered verification

Begin with access logs and capture the full User-Agent, source IP, ASN, reverse DNS, path, method, status, response size, redirects, timestamp, and request rate. A token alone is not authentication. A request claiming kagi-fetcher/1.0 may be a genuine product request, an old registry pattern, or an unrelated client copying the string.

Compare the observed token with Kagi's official documentation. If the request instead declares Kagibot, use Kagi's published kagi.com/bot page and its source IPs for that separate identity. If it declares kagi-fetcher, do not silently promote a KagiBot match into verification.

Evaluate /robots.txt and page directives independently. For a page that should not be indexed, you may also use:

configuration / code
<meta name="robots" content="noindex, nofollow">
configuration / code
X-Robots-Tag: noindex, nofollow

These signals do not establish that a user-triggered fetcher will stop reading the response, and they do not protect private content. Use route authorization, signed URLs, and rate limits where confidentiality or cost matters.

WAF and Nginx remediation examples

If you have confirmed unwanted requests carrying the token, a narrow WAF rule can block the declared identity:

configuration / code
{
  "description": "Block observed kagi-fetcher token",
  "expression": "lower(http.user_agent) contains \"kagi-fetcher\"",
  "action": "block"
}

For Nginx, scope enforcement to routes that should not be fetched by an unverified assistant token:

configuration / code
map $http_user_agent $block_kagi_fetcher {
    default 0;
    ~*kagi-fetcher 1;
}

server {
    location ~ ^/(private|account|checkout|internal|api)/ {
        if ($block_kagi_fetcher) { return 403; }
        try_files $uri $uri/ =404;
    }
}

Test the rule against Kagi Search's documented Kagibot, ordinary browsers, internal health checks, and permitted assistant workflows. A header match is easy to spoof and may miss a request made through another Kagi component. If you need reliable protection, combine edge rules with authentication and behavioral controls.

Review checklist

Search logs for the exact kagi-fetcher/1.0 token and separate it from Kagibot. Record request volume, paths, response sizes, source networks, and whether the traffic appears user-triggered or systematic. Review Kagi's current AI documentation and Kagibot page without treating either as a kagi-fetcher contract.

Decide whether your goal is to prevent private page fetching, preserve search visibility, protect bandwidth, or restrict data extraction. Publish a targeted robots group if it communicates the desired preference, then enforce high-risk routes using WAF, Nginx, authentication, and rate limits. Re-test public docs, private pages, APIs, and search-crawler traffic independently. Revisit this profile if Kagi publishes a dedicated kagi-fetcher specification.

References

  1. Kagi About KagiBot — official documentation for Kagi's separate search crawler, including its User-Agent, IPs, and robots behavior.
  2. Kagi AI — official documentation for Kagi's AI product features; it does not publish a separate kagi-fetcher contract.
  3. Google Robots.txt Introduction — general explanation of crawler directives and their limitations.

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.