Bot directory / ai-training

KendraBot: Robots.txt & Crawl Policy Reference

Technical reference for the KendraBot registry label and Amazon Kendra Web Crawler. Learn why the documented AWS token is amazon-kendra and how to verify actual traffic.

AI Summary: Amazon Kendra Web Crawler is an enterprise search connector that customers configure to index selected websites. AWS's current robots.txt examples use the token amazon-kendra, not KendraBot; AWS also states that Kendra is no longer open to new customers. Treat KendraBot as a registry label requiring log verification, and apply AWS's documented policy only when the exact amazon-kendra token is observed.

Role and policy boundary

Amazon Kendra is an enterprise search service that indexes documents selected by an AWS customer so users can search that organization's content. AWS documentation says the Web Crawler can crawl public HTTPS websites and, in supported configurations, internal websites through a proxy or authentication. AWS also states that customers must obtain authorization before indexing a website and that aggressive crawling of sites the customer does not own is not acceptable use.

The registry entry uses the display label KendraBot, but AWS's robots documentation publishes amazon-kendra in its examples. This is a material identity distinction. There is no basis to claim that every request declaring KendraBot is an official Amazon Kendra request, and it would be unsafe to copy a rule for amazon-kendra into a policy decision without checking live logs.

AWS documents Allow and Disallow handling for the amazon-kendra token. To allow Kendra to crawl a public documentation area while excluding sensitive routes:

configuration / code
User-agent: amazon-kendra
Allow: /docs/
Allow: /public-reference/
Disallow: /credential-pages/
Disallow: /account/
Disallow: /internal/

To block the current documented token:

configuration / code
User-agent: amazon-kendra
Disallow: /

Do not assume that a KendraBot rule controls amazon-kendra, or vice versa. Kendra indexing is customer-configured enterprise search, not evidence that AWS uses your content to train a general-purpose model.

Layered verification

Begin with access logs and search separately for KendraBot and amazon-kendra. Record the full User-Agent, source IP, ASN, reverse DNS, path, method, status, response size, redirect chain, timestamp, and request rate. A self-declared token is not identity authentication.

Compare the observed token with the AWS documentation. If the request uses amazon-kendra, evaluate it against the documented robots examples and the configured site scope. If it uses KendraBot, mark the attribution as unresolved until a current source confirms it. Do not use AWS's service policy to authenticate a different label.

Evaluate /robots.txt independently. Confirm the canonical host, status, content type, exact user-agent group, and path match. Kendra's web connector can use authenticated or proxied sources, so public robots behavior does not provide access to private data. For page-level discoverability, inspect:

configuration / code
<meta name="robots" content="noindex, nofollow">
configuration / code
X-Robots-Tag: noindex, nofollow

These directives are separate from authentication and do not protect enterprise-only routes. Use authorization, network segmentation, and route controls for confidential content.

WAF and Nginx remediation examples

If logs confirm unwanted requests carrying the exact registry label, a narrow WAF rule can block the claim:

configuration / code
{
  "description": "Block observed KendraBot token",
  "expression": "lower(http.user_agent) contains \"kendrabot\"",
  "action": "block"
}

If the goal is to block the token AWS currently documents, use a separate rule rather than silently merging identities:

configuration / code
{
  "description": "Block documented Amazon Kendra token",
  "expression": "lower(http.user_agent) contains \"amazon-kendra\"",
  "action": "block"
}

For Nginx, scope a registry-label block to sensitive paths while you investigate:

configuration / code
map $http_user_agent $block_kendra_label {
    default 0;
    ~*KendraBot 1;
}

server {
    location ~ ^/(internal|account|credential-pages|api)/ {
        if ($block_kendra_label) { return 403; }
        try_files $uri $uri/ =404;
    }
}

A header match can be spoofed and can miss a Kendra connector using amazon-kendra. Test the intended token independently, and protect private enterprise content with authentication rather than relying on WAF or robots rules alone.

Review checklist

Search logs for both KendraBot and amazon-kendra; never treat them as synonyms. Confirm the requested paths, status codes, response sizes, source network, and rate. Read AWS's current Web Crawler documentation and note that Amazon Kendra is no longer open to new customers; similar new workloads may use Amazon Bedrock Knowledge Bases instead.

Decide whether your goal is to preserve enterprise-search availability, block a stale registry token, protect private content, or reduce crawl load. Publish a specific robots group for the exact observed token, enforce high-risk routes with authentication and WAF controls, and test public docs, credential pages, APIs, and authenticated paths separately. Revisit this profile if AWS publishes a current KendraBot identity or if the connector token changes.

References

  1. Configuring robots.txt for Amazon Kendra Web Crawler — official AWS token and robots behavior; examples use amazon-kendra.
  2. Amazon Kendra Web Crawler — official scope, authorization, authentication, connector, and availability notes.
  3. Google Robots.txt Introduction — general explanation of crawler directives and their limitations.

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.