Amzn-SearchBot: Robots.txt & Crawl Policy Reference
Comprehensive guide to Amzn-SearchBot, an Amazon web crawler. Understand its role in Amazon search experiences and how it differs from Amazonbot.
AI Summary:
Amzn-SearchBotis a web crawler operated by Amazon, primarily used to index content to improve search experiences within Amazon products and services. Crucially, Amazon explicitly states that this specific bot does not crawl for AI model training, distinguishing it from the broaderAmazonbot.
Role and policy boundary
Amzn-SearchBot is designed to gather data to enhance Amazon's search capabilities, ensuring your content is discoverable and can be cited in Amazon's search ecosystem.
This crawler has a specific policy boundary: Amazon documents that Amzn-SearchBot (along with Amzn-User) is not used to collect data for training foundational AI models (like AWS Bedrock models). If your goal is to opt out of AI training but remain visible in Amazon's search features, you can block Amazonbot while allowing Amzn-SearchBot.
A robots rule is a declaration of intent; it does not replace authentication, authorization, or rate limiting. Start with a dedicated group:
User-agent: Amzn-SearchBot
Allow: /
Disallow: /staging/
Disallow: /internal/
To stop access for the entire site, use:
User-agent: Amzn-SearchBot
Disallow: /
Avoid assuming that User-agent: * expresses the same business intent. A wildcard can affect assistant and training crawlers too, and it makes later audits harder because the source of the decision is less specific.
Layered verification
Verify the same URL through each control plane instead of assuming that one green signal represents the whole request path. Compare the bot-specific robots group, the page-level metadata, and the response headers captured at the public edge.
Amazon officially states that Amzn-SearchBot respects robots.txt directives. However, like the broader Amazonbot, it historically has struggled with or ignored directives like crawl-delay and page-level meta tags (noindex, nofollow). To reliably block it, a strict Disallow in robots.txt is required, and server-level enforcement may be necessary if it crawls too aggressively.
The Policy Engine evaluates the selected user-agent, path scope, and the other supplied layers independently. It can therefore explain why a bot is allowed while another is blocked, rather than returning one blended website score.
Page-level directives can still override the intended outcome for indexing:
<meta name="robots" content="noindex, nofollow">
X-Robots-Tag: noindex, nofollow
If a response uses these tags, the report marks the result as blocked or conflicting even if the crawler-specific robots group is permissive. This is especially important for canonical pages served through an edge cache where headers may differ from the origin response.
WAF and Nginx remediation examples
If you find that Amzn-SearchBot is ignoring robots.txt (as some webmasters have reported on Cloudflare forums) or consuming excessive bandwidth, you should block it at the WAF level:
{
"description": "Block Amazon SearchBot",
"expression": "lower(http.user_agent) contains \"amzn-searchbot\"",
"action": "block"
}
Use your platform's actual middleware response pattern rather than copying this simplified example without review. Never place a secret, verification token, or internal policy identifier in a public header.
map $http_user_agent $block_amzn_searchbot_private {
default 0;
~*Amzn-SearchBot 1;
}
server {
location ~ ^/(admin|account|private|licensed|internal|api)/ {
if ($block_amzn_searchbot_private) { return 403; }
try_files $uri $uri/ =404;
}
}
Review checklist
Use this checklist after every policy change and after a CDN or WAF migration. Record the request URL, User-Agent, HTTP status, final redirect, and the exact evidence used to reach the decision.
Verify that the dedicated group appears before relying on a wildcard, test a representative public and private path, and compare live response headers with robots.txt. Keep the policy close to the content owner's intent and record whether the site wants discovery, citation, or no access at all.
Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.