PerplexityBot — Perplexity search crawler
A technical reference for PerplexityBot search discovery controls.
AI Summary: PerplexityBot is a search-oriented crawler. Keep its policy block explicit if the site wants to reason about discovery separately from training.
Role and policy boundary
Treat this User-Agent as a distinct policy identity. The correct decision depends on the bot purpose, the paths it requests, and the response controls applied by the origin, CDN, and application layer. A robots rule is a declaration of intent; it does not replace authentication, authorization, or rate limiting.
PerplexityBot is included in the search group. A site that wants its public articles to be discoverable can state that preference explicitly:
User-agent: PerplexityBot
Allow: /articles/
Allow: /reference/
Disallow: /preview/
Disallow: /drafts/
If the site does not want search retrieval, the most direct policy is:
User-agent: PerplexityBot
Disallow: /
Use explicit paths when only a subset of the site is ready. This makes the policy reviewable and reduces the chance that a new private route accidentally inherits a broad allow rule. For sensitive content, add real authentication and authorization; robots.txt is not a security boundary.
Layered verification
Verify the same URL through each control plane instead of assuming that one green signal represents the whole request path. Compare the bot-specific robots group, the page-level metadata, and the response headers captured at the public edge.
The catalog intentionally separates PerplexityBot from GPTBot and Bytespider. A site can permit search discovery while blocking training-oriented agents:
User-agent: PerplexityBot
Allow: /
User-agent: GPTBot
Disallow: /
User-agent: Bytespider
Disallow: /
The live scan compares this robots evidence with page and response directives. For example, a noindex header on /articles/launch creates a finding even when PerplexityBot is allowed in robots.txt. The report should surface both the permissive and restrictive sources instead of hiding the conflict behind a single score.
WAF and Nginx remediation examples
A Cloudflare Custom Rule can start in log mode:
{
"description": "Audit PerplexityBot search access",
"expression": "lower(http.user_agent) contains \"perplexitybot\"",
"action": "log",
"note": "Move to a scoped block only after reviewing public search requirements."
}
A simple Nginx map can classify the crawler for an existing request policy:
map $http_user_agent $perplexitybot_policy {
default allow;
~*"PerplexityBot" allow;
}
add_header X-AIBotCheck-Policy $perplexitybot_policy always;
Do not expose internal reasoning, secrets, or correlation identifiers in headers. If your organization does not want a diagnostic header in production, keep classification entirely inside the edge rule.
Review checklist
Use this checklist after every policy change and after a CDN or WAF migration. Record the request URL, User-Agent, HTTP status, final redirect, and the exact evidence used to reach the decision.
Fetch the canonical URL and /robots.txt, confirm the final URL after redirects, and inspect whether the content type is text. Compare a public article, a preview URL, and a private route. If the site uses a CDN, repeat the check from the intended production hostname because origin and edge headers can differ. Store the raw evidence with a timestamp in your own audit system rather than assuming the current response will remain unchanged.
Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.