Bot directory / ai-training

apifybot: Robots.txt & Crawl Policy Reference

Comprehensive guide to apifybot, the default crawler for the Apify platform. Understand its role in web scraping, data extraction, and automation.

AI Summary: apifybot (often seen as ApifyBot) is associated with Apify, a popular platform for web scraping, data extraction, and web automation. It is frequently used by developers to build datasets, which can include data for AI training, LLM applications, or RAG pipelines.

Role and policy boundary

Apify provides tools (like Crawlee) for users to create their own scrapers (called Actors). The default User-Agent for many of these operations is apifybot.

Because Apify is a generalized scraping platform, the data collected by apifybot could be used for anything from price monitoring to training large language models. If your policy is to prevent unauthorized third-party scraping and data extraction, you should block this bot.

A robots rule is a declaration of intent; it does not replace authentication, authorization, or rate limiting. Start with a dedicated group:

configuration / code
User-agent: apifybot
Allow: /
Disallow: /staging/
Disallow: /internal/

To stop access for the entire site, use:

configuration / code
User-agent: apifybot
Disallow: /

Avoid assuming that User-agent: * expresses the same business intent. A wildcard can affect assistant and training crawlers too, and it makes later audits harder because the source of the decision is less specific.

Layered verification

Verify the same URL through each control plane instead of assuming that one green signal represents the whole request path. Compare the bot-specific robots group, the page-level metadata, and the response headers captured at the public edge.

Apify generally encourages its users to respect robots.txt, and the default apifybot User-Agent can be blocked via standard directives. However, because Apify is a platform used by third parties, enforcement ultimately relies on the individual user's configuration. Users can easily configure their scrapers to spoof User-Agents (e.g., pretending to be a normal Chrome browser) or use residential proxies, making detection more complex than a single bot signature.

The Policy Engine evaluates the selected user-agent, path scope, and the other supplied layers independently. It can therefore explain why a bot is allowed while another is blocked, rather than returning one blended website score.

Page-level directives can still override the intended outcome for indexing:

configuration / code
<meta name="robots" content="noindex, nofollow">
configuration / code
X-Robots-Tag: noindex, nofollow

If a response uses these tags, the report marks the result as blocked or conflicting even if the crawler-specific robots group is permissive. This is especially important for canonical pages served through an edge cache where headers may differ from the origin response.

WAF and Nginx remediation examples

Blocking the apifybot User-Agent is a good first step. However, defending against platform-based scraping often requires more advanced bot management solutions (like rate limiting, behavioral analysis, or IP reputation scoring) if users spoof their User-Agents.

To block the default Apify bot at the network layer:

configuration / code
{
  "description": "Block Apify default crawler",
  "expression": "lower(http.user_agent) contains \"apifybot\"",
  "action": "block"
}

Use your platform's actual middleware response pattern rather than copying this simplified example without review. Never place a secret, verification token, or internal policy identifier in a public header.

configuration / code
map $http_user_agent $block_apifybot_private {
    default 0;
    ~*apifybot 1;
}

server {
    location ~ ^/(admin|account|private|licensed|internal|api)/ {
        if ($block_apifybot_private) { return 403; }
        try_files $uri $uri/ =404;
    }
}

Review checklist

Use this checklist after every policy change and after a CDN or WAF migration. Record the request URL, User-Agent, HTTP status, final redirect, and the exact evidence used to reach the decision.

Verify that the dedicated group appears before relying on a wildcard, test a representative public and private path, and compare live response headers with robots.txt. Keep the policy close to the content owner's intent and record whether the site wants discovery, citation, or no access at all.


Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.