Glossary / 21 terms

AI Bot Glossary

Technical definitions and architecture guides for key concepts across AI crawlers, robots.txt protocols, edge verification, and Generative Engine Optimization.

AI Search Crawler

A specialized web crawler operated by an AI search engine to retrieve real-time web content for live retrieval-augmented generation (RAG) and citations.

ai-search-crawlerRead definition →

AI Training Crawler

A web crawler deployed by AI laboratories to extract large text and multimodal datasets from the open web for pre-training and fine-tuning foundational models.

ai-training-crawlerRead definition →

Assistant Crawler

An on-demand fetcher triggered directly by an end-user interacting with an AI chat interface or custom agent to browse a specific URL in real time.

assistant-crawlerRead definition →

Bot Management

The architectural practice of detecting, classifying, and mitigating automated traffic across edge networks, application firewalls, and server configurations.

bot-managementRead definition →

Generative AI

Artificial intelligence systems capable of generating novel text, code, imagery, or synthesis by predicting probability distributions learned from training data.

generative-aiRead definition →

Headless Browser

A web browser running without a graphical user interface (GUI) controlled via automated scripts to render JavaScript, take screenshots, and scrape dynamic applications.

headless-browserRead definition →

Web Indexing

The computational process of parsing, organizing, and storing web documents into an inverted index database to enable rapid query retrieval.

IP Allowlist

A network security control that restricts incoming traffic strictly to explicitly approved IP addresses and CIDR prefixes, rejecting all unlisted connections.

Meta Robots Tag

An HTML element placed in the document head that provides page-level instructions to search engine crawlers regarding indexing, link following, and snippet rendering.

Noindex Directive

A robots directive used in HTML meta tags or HTTP headers that explicitly commands search engines and crawlers not to index the designated page.

Rate Limiting

A network traffic management strategy that restricts the number of requests a single client or IP address can submit to a server within a defined timeframe.

rate-limitingRead definition →

Robots.txt Protocol

A plain text configuration file placed at the root of a domain that establishes declarative access policies for automated web crawlers according to RFC 9309.

Schema.org & JSON-LD

The shared structured-data vocabulary used to declare entities, authorship, and page meaning to search engines and AI crawlers via JSON-LD markup.

User-Agent Spoofing

The practice of fabricating or mimicking the User-Agent header of legitimate search engines or AI assistants to bypass access controls, scrapers, and security filters.

User-Agent Header

An HTTP request header that identifies the client application, operating system, vendor, and version submitting the network request to a web server.

Web Scraping

The programmatic extraction of structured data, text, and media assets from websites by automated bots for aggregation, analysis, or machine learning ingestion.