Amzn-User: Robots.txt & Crawl Policy Reference
Comprehensive guide to Amzn-User, Amazon's user-triggered fetcher. Learn how it operates on behalf of users and how to manage access.
AI Summary:
Amzn-Useris an on-demand, user-triggered fetcher operated by Amazon. It retrieves web pages when an Amazon product feature or user explicitly requests external content. Importantly, Amazon documents that it is not used to crawl for AI model training.
Role and policy boundary
Unlike Amazon's background crawlers (Amazonbot and Amzn-SearchBot), Amzn-User acts as a live browsing agent on behalf of a user's request. It is not used for bulk background data collection or foundational model training (such as for AWS Bedrock).
Allowing Amzn-User ensures that if an Amazon user requests information from your site through an Amazon service (like an Alexa query that needs to fetch a specific URL live), the request can succeed. Blocking it means those specific user-initiated features will fail to read your content.
A robots rule is a declaration of intent; it does not replace authentication, authorization, or rate limiting. Start with a dedicated group:
User-agent: Amzn-User
Allow: /
Disallow: /staging/
Disallow: /internal/
To stop access for the entire site, use:
User-agent: Amzn-User
Disallow: /
Avoid assuming that User-agent: * expresses the same business intent. A wildcard can affect assistant and training crawlers too, and it makes later audits harder because the source of the decision is less specific.
Layered verification
Verify the same URL through each control plane instead of assuming that one green signal represents the whole request path. Compare the bot-specific robots group, the page-level metadata, and the response headers captured at the public edge.
Because it operates on behalf of a user, Amzn-User may have different robots.txt compliance characteristics compared to background crawlers, depending on the specific Amazon product initiating the fetch. However, Amazon officially states that their crawlers look for robots.txt at the host level, so targeting the User-Agent is the standard approach.
The Policy Engine evaluates the selected user-agent, path scope, and the other supplied layers independently. It can therefore explain why a bot is allowed while another is blocked, rather than returning one blended website score.
Page-level directives can still override the intended outcome for indexing:
<meta name="robots" content="noindex, nofollow">
X-Robots-Tag: noindex, nofollow
If a response uses these tags, the report marks the result as blocked or conflicting even if the crawler-specific robots group is permissive. This is especially important for canonical pages served through an edge cache where headers may differ from the origin response.
WAF and Nginx remediation examples
To prevent Amazon's user-triggered features from accessing your site (for example, if you want to completely block all Amazon infrastructure regardless of user intent), block the User-Agent at the network layer:
{
"description": "Block Amazon User Fetcher",
"expression": "lower(http.user_agent) contains \"amzn-user\"",
"action": "block"
}
Use your platform's actual middleware response pattern rather than copying this simplified example without review. Never place a secret, verification token, or internal policy identifier in a public header.
map $http_user_agent $block_amzn_user_private {
default 0;
~*Amzn-User 1;
}
server {
location ~ ^/(admin|account|private|licensed|internal|api)/ {
if ($block_amzn_user_private) { return 403; }
try_files $uri $uri/ =404;
}
}
Review checklist
Use this checklist after every policy change and after a CDN or WAF migration. Record the request URL, User-Agent, HTTP status, final redirect, and the exact evidence used to reach the decision.
Verify that the dedicated group appears before relying on a wildcard, test a representative public and private path, and compare live response headers with robots.txt. Keep the policy close to the content owner's intent and record whether the site wants discovery, citation, or no access at all.
Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.