Policy Reference

Access control layering.

Guidance on governing search, assistant, and training crawlers using explicit policies and layered edge verification.

Optimizing for Google AI Overviews

A technical architecture guide to ensuring web content is indexed, synthesized, and cited within Google AI Overviews and Search Generative Experience.

Read policy →

Apache Web Server: Blocking AI Crawlers & Scrapers

Production-ready Apache HTTP Server (.htaccess and httpd.conf) configurations to block, throttle, or isolate aggressive AI crawlers and automated bots.

Read policy →

Assistant Access Policy: Managing User-Driven AI Bots

A guide to managing web crawlers triggered directly by users interacting with AI assistants, ensuring they can retrieve public information without compromising security.

Read policy →

Balancing Traditional SEO with AI Training Opt-Outs

A strategic framework for protecting intellectual property from foundational AI training without sacrificing organic search rankings or AI search citations.

Read policy →

ChatGPT Search Visibility & Crawler Optimization

A technical guide to configuring your web infrastructure for citation and discovery in ChatGPT Search and OpenAI conversational answers.

Read policy →

Cloudflare WAF: Managing AI Crawlers at the Edge

Enterprise guide to configuring Cloudflare WAF Custom Rules, Bot Management, and AI Crawlers protection to govern automated bot access.

Read policy →

Copyright, Fair Use, and AI Model Training

An analysis of legal frameworks, intellectual property boundaries, and copyright opt-out mechanisms for AI training datasets.

Read policy →

DMCA Takedowns & Copyright Notices for AI Training Datasets

A procedural guide for publishers issuing DMCA notices and copyright claims against AI labs and training dataset repositories.

Read policy →

The EU AI Act: Web Scraping, Transparency & Publisher Rights

A regulatory analysis of the European Union AI Act, copyright compliance mandates, and transparency obligations for general-purpose AI providers.

Read policy →

Identifying Spoofed AI Bots: Verification & Forensics

A technical forensics guide to detecting and blocking malicious scrapers that forge User-Agent headers to masquerade as legitimate AI crawlers.

Read policy →

IndexNow Protocol: Real-Time Indexation for AI & Search

How to implement the IndexNow protocol to instantly notify Microsoft Bing, Yandex, and AI search engines of content changes.

Read policy →

The /llms.txt Standard: Documentation for AI Systems

A complete architectural guide to creating, structuring, and serving an /llms.txt file to guide AI models, crawlers, and coding agents.

Read policy →

Nginx Configuration: Blocking & Throttling AI Bots

High-performance Nginx map directives, rate limits, and status codes to block, rate-limit, and segregate AI crawlers at the edge.

Read policy →

Perplexity AI: Citation Guidelines & Crawler Optimization

Technical and architectural specifications for ensuring your website is crawled by PerplexityBot and cited in Perplexity answers.

Read policy →

Rate-Limiting Strategies for Automated AI Crawlers

Architectural patterns for throttling, queueing, and governing aggressive automated crawlers to safeguard origin stability.

Read policy →

Robots Layering Policy: Coordinating Access Controls

A technical guide to understanding how robots.txt, page-level metadata, HTTP headers, and WAF rules interact to control AI bot access.

Read policy →

Robots.txt Wildcards & Pattern Matching (RFC 9309)

A deep-dive technical guide to pattern matching rules, wildcards (*), and end-of-string anchors ($) in RFC 9309 robots.txt files.

Read policy →

Search Visibility Policy: Optimizing for AI Search Engines

A guide to configuring your website to be discovered and indexed by AI-powered search engines and answer engines while maintaining control over your content.

Read policy →

Drafting Terms of Service Clauses to Restrict AI Scraping

A practical guide to drafting enforceable Terms of Service provisions that prohibit automated scraping, AI model training, and data commercialization.

Read policy →

Training Opt-Out Policy: How to Block AI Data Scraping

A comprehensive guide to blocking AI crawlers from using your website's content for training large language models (LLMs) while maintaining search visibility.

Read policy →