← Bot Directory/PerplexityBot
Bot directory / AI Search Bot

PerplexityBot — Perplexity search crawler

A technical reference for PerplexityBot search discovery controls.

AI Summary: PerplexityBot is a search-oriented crawler. Keep its policy block explicit if the site wants to reason about discovery separately from training.

Role and policy boundary

Treat this User-Agent as a distinct policy identity. The correct decision depends on the bot purpose, the paths it requests, and the response controls applied by the origin, CDN, and application layer. A robots rule is a declaration of intent; it does not replace authentication, authorization, or rate limiting.

PerplexityBot is included in the search group. A site that wants its public articles to be discoverable can state that preference explicitly:

configuration / code
User-agent: PerplexityBot
Allow: /articles/
Allow: /reference/
Disallow: /preview/
Disallow: /drafts/

If the site does not want search retrieval, the most direct policy is:

configuration / code
User-agent: PerplexityBot
Disallow: /

Use explicit paths when only a subset of the site is ready. This makes the policy reviewable and reduces the chance that a new private route accidentally inherits a broad allow rule. For sensitive content, add real authentication and authorization; robots.txt is not a security boundary.

Layered verification

Verify the same URL through each control plane instead of assuming that one green signal represents the whole request path. Compare the bot-specific robots group, the page-level metadata, and the response headers captured at the public edge.

The catalog intentionally separates PerplexityBot from GPTBot and Bytespider. A site can permit search discovery while blocking training-oriented agents:

configuration / code
User-agent: PerplexityBot
Allow: /

User-agent: GPTBot
Disallow: /

User-agent: Bytespider
Disallow: /

The live scan compares this robots evidence with page and response directives. For example, a noindex header on /articles/launch creates a finding even when PerplexityBot is allowed in robots.txt. The report should surface both the permissive and restrictive sources instead of hiding the conflict behind a single score.

WAF and Nginx remediation examples

A Cloudflare Custom Rule can start in log mode:

configuration / code
{
  "description": "Audit PerplexityBot search access",
  "expression": "lower(http.user_agent) contains \"perplexitybot\"",
  "action": "log",
  "note": "Move to a scoped block only after reviewing public search requirements."
}

A simple Nginx map can classify the crawler for an existing request policy:

configuration / code
map $http_user_agent $perplexitybot_policy {
    default allow;
    ~*"PerplexityBot" allow;
}

add_header X-AIBotCheck-Policy $perplexitybot_policy always;

Do not expose internal reasoning, secrets, or correlation identifiers in headers. If your organization does not want a diagnostic header in production, keep classification entirely inside the edge rule.

Review checklist

Use this checklist after every policy change and after a CDN or WAF migration. Record the request URL, User-Agent, HTTP status, final redirect, and the exact evidence used to reach the decision.

Fetch the canonical URL and /robots.txt, confirm the final URL after redirects, and inspect whether the content type is text. Compare a public article, a preview URL, and a private route. If the site uses a CDN, repeat the check from the intended production hostname because origin and edge headers can differ. Store the raw evidence with a timestamp in your own audit system rather than assuming the current response will remain unchanged.


Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.