Assistant Access Policy: Managing User-Driven AI Bots
A guide to managing web crawlers triggered directly by users interacting with AI assistants, ensuring they can retrieve public information without compromising security.
AI Summary: Assistant access refers to web requests made by AI chatbots (like ChatGPT or Claude) on behalf of a user in real-time. These bots fetch specific URLs provided by the user to summarize or analyze content. Webmasters should allow these user-agent tokens to access public pages to enhance user experience, but must strictly enforce authentication at the server level to prevent unauthorized access to private data.
The Nature of Assistant Crawlers
Unlike traditional search engine crawlers that systematically index the entire web, or training crawlers that scrape massive datasets in the background, assistant crawlers are user-driven.
When a user pastes a URL into ChatGPT and asks, "Summarize this article," OpenAI's servers dispatch a request to that URL using the ChatGPT-User user-agent. This is a targeted, synchronous request designed to retrieve the content of that specific page for the user's immediate benefit.
Common assistant user-agents include:
ChatGPT-User(OpenAI / ChatGPT)Claude-User(Anthropic / Claude)Perplexity-User(Perplexity AI)MistralAI-User(Mistral AI)
Blocking these agents degrades the experience for your users, preventing them from using their preferred AI tools to interact with your public content.
Recommended Layering for Assistant Access
Because assistant requests are triggered by users, they should generally be treated similarly to standard web browsers when accessing public content. However, because they are automated scripts, they must be explicitly managed in robots.txt and strictly bounded by server security.
Step 1: Configure robots.txt
Use robots.txt to explicitly allow assistant bots to access your public content, while disallowing them from areas where automated retrieval is inappropriate.
# Allow ChatGPT to fetch public pages for users
User-agent: ChatGPT-User
Allow: /
Disallow: /admin/
Disallow: /private/
Disallow: /api/
# Allow Claude to fetch public pages for users
User-agent: Claude-User
Allow: /
Disallow: /admin/
Disallow: /private/
Disallow: /api/
Step 2: Enforce Security Boundaries
The most critical aspect of an assistant access policy is recognizing that AI assistants do not share the user's browser session.
If a user is logged into your website and asks ChatGPT to summarize a private page (e.g., https://yoursite.com/my-account/billing), ChatGPT will attempt to fetch that URL. Because the request comes from OpenAI's servers—not the user's browser—it will not have the user's authentication cookies.
Your server must return a 401 Unauthorized or 403 Forbidden status, or redirect to a login page. If your server relies on client-side JavaScript to enforce security (a major vulnerability), the AI bot might scrape the raw HTML or JSON data, exposing private information.
WAF and Nginx Remediation Examples
You should not block assistant bots globally, but you may want to monitor them or explicitly block them from sensitive API endpoints to prevent abuse.
WAF JSON Example
Monitor assistant bot traffic to ensure it is not being used for high-volume scraping:
{
"description": "Monitor user-driven AI assistant requests",
"expression": "lower(http.user_agent) contains \"chatgpt-user\" or lower(http.user_agent) contains \"claude-user\"",
"action": "log"
}
Nginx Remediation Example
Ensure that your server configuration strictly requires authentication for private routes, regardless of the user-agent. You can also explicitly reject assistant bots from API routes where they have no legitimate purpose:
map $http_user_agent $is_assistant_bot {
default 0;
~*ChatGPT-User 1;
~*Claude-User 1;
~*Perplexity-User 1;
}
server {
# Public routes: Allow assistant bots
location / {
try_files $uri $uri/ /index.html;
}
# API routes: Block assistant bots, as they should only fetch HTML pages
location /api/ {
if ($is_assistant_bot) {
return 403;
}
# Standard API configuration follows...
}
# Private routes: Require strict authentication
location ~ ^/(admin|account|private)/ {
# If the request lacks a valid session/token, reject it.
# This naturally protects against assistant bots lacking cookies.
auth_request /auth-verify;
error_page 401 = /login;
}
}
Review Checklist
When implementing an assistant access policy, verify the following:
- [ ] Have you explicitly allowed user-driven bots (e.g.,
ChatGPT-User,Claude-User) in yourrobots.txtto access public content? - [ ] Are sensitive, administrative, or API routes explicitly disallowed in
robots.txt? - [ ] Crucial: Does your server correctly return a
401,403, or redirect to login when an unauthenticated request (like one from an AI server) attempts to access a private URL? - [ ] Have you verified that your site does not rely solely on client-side JavaScript (e.g., hiding DOM elements) to protect sensitive data?
- [ ] Are you monitoring log files to ensure assistant bots are not being weaponized for high-volume scraping?
Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.