← Bot Directory/ApifyWebsiteContentCrawler
Bot directory / ai-training

ApifyWebsiteContentCrawler: Robots.txt & Crawl Policy Reference

Comprehensive guide to ApifyWebsiteContentCrawler, an Apify Actor used to extract text and markdown for AI and LLM pipelines.

AI Summary: ApifyWebsiteContentCrawler is a specific User-Agent associated with the "Website Content Crawler" Actor on the Apify platform. It is explicitly designed to perform deep crawls of websites to extract clean text, HTML, and Markdown, primarily to feed AI models, LLM applications, vector databases, and RAG pipelines.

Role and policy boundary

Unlike the generic apifybot, this crawler is highly specialized for AI data ingestion. It is a tool used by third-party developers and companies to scrape your website's content for their own machine learning pipelines.

If your policy is to prevent unauthorized third-party scraping for AI training or RAG (Retrieval-Augmented Generation) applications, you must explicitly block this crawler. Allowing it means you are permitting users of the Apify platform to easily convert your website into a dataset.

A robots rule is a declaration of intent; it does not replace authentication, authorization, or rate limiting. Start with a dedicated group:

configuration / code
User-agent: ApifyWebsiteContentCrawler
Allow: /
Disallow: /staging/
Disallow: /internal/

To stop access for the entire site, use:

configuration / code
User-agent: ApifyWebsiteContentCrawler
Disallow: /

Avoid assuming that User-agent: * expresses the same business intent. A wildcard can affect assistant and training crawlers too, and it makes later audits harder because the source of the decision is less specific.

Layered verification

Verify the same URL through each control plane instead of assuming that one green signal represents the whole request path. Compare the bot-specific robots group, the page-level metadata, and the response headers captured at the public edge.

While the default Apify Actor respects robots.txt when using this User-Agent, webmasters should be aware that Apify users can easily modify the configuration of their scrapers to ignore robots.txt or spoof their User-Agent to look like a standard web browser. Therefore, relying solely on robots.txt is often a weak defense against determined scrapers using this platform.

The Policy Engine evaluates the selected user-agent, path scope, and the other supplied layers independently. It can therefore explain why a bot is allowed while another is blocked, rather than returning one blended website score.

Page-level directives can still override the intended outcome for indexing:

configuration / code
<meta name="robots" content="noindex, noai">
configuration / code
X-Robots-Tag: noindex, noai

If a response uses these tags, the report marks the result as blocked or conflicting even if the crawler-specific robots group is permissive. This is especially important for canonical pages served through an edge cache where headers may differ from the origin response.

WAF and Nginx remediation examples

To effectively block the Apify Website Content Crawler, you should start by blocking its declared User-Agent at the WAF level. However, for robust protection against Apify scrapers, you should also implement rate limiting and behavioral bot management.

configuration / code
{
  "description": "Block Apify Website Content Crawler",
  "expression": "lower(http.user_agent) contains \"apifywebsitecontentcrawler\"",
  "action": "block"
}

Use your platform's actual middleware response pattern rather than copying this simplified example without review. Never place a secret, verification token, or internal policy identifier in a public header.

configuration / code
map $http_user_agent $block_apifywebsitecontentcrawler_private {
    default 0;
    ~*ApifyWebsiteContentCrawler 1;
}

server {
    location ~ ^/(admin|account|private|licensed|internal|api)/ {
        if ($block_apifywebsitecontentcrawler_private) { return 403; }
        try_files $uri $uri/ =404;
    }
}

Review checklist

Use this checklist after every policy change and after a CDN or WAF migration. Record the request URL, User-Agent, HTTP status, final redirect, and the exact evidence used to reach the decision.

Verify that the dedicated group appears before relying on a wildcard, test a representative public and private path, and compare live response headers with robots.txt. Keep the policy close to the content owner's intent and record whether the site wants discovery, citation, or no access at all.


Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.