← Bot Directory/AppEngine-Google
Bot directory / scraper

AppEngine-Google: Robots.txt & Crawl Policy Reference

Technical reference for AppEngine-Google. Learn why this User-Agent represents third-party applications on Google Cloud, not an official Google crawler.

AI Summary: AppEngine-Google is not an official Google crawler (like Googlebot). It is the default User-Agent appended to outbound HTTP requests made by third-party applications hosted on Google App Engine. Because this traffic originates from arbitrary developers using Google's cloud infrastructure, it does not inherently respect robots.txt and its purpose varies entirely by the application. Treat it as untrusted third-party traffic.

Role and policy boundary

The registry associates AppEngine-Google with Google App Engine. Historically, requests made using App Engine's URL Fetch service or default HTTP libraries include AppEngine-Google; (+http://code.google.com/appengine; appid: [APP_ID]) in the User-Agent header.

It is critical to understand that this traffic does not represent Google. It represents a developer who deployed code to Google Cloud Platform. The purpose could be anything: a legitimate RSS feed reader, a webhook delivery system, a malicious scraper, or an AI data extraction tool.

Because the crawler is not a centralized Google service, it does not automatically respect robots.txt. Whether it honors directives depends entirely on whether the individual developer programmed their application to parse and obey robots.txt.

If you want to communicate a restriction, you can publish a rule, but expect it to be largely ignored by aggressive scrapers:

configuration / code
User-agent: AppEngine-Google
Disallow: /

Robots.txt is advisory and cannot protect private or licensed content. Use authentication, authorization, signed URLs, and origin controls for those boundaries.

Layered verification

Start with raw access logs and preserve the complete User-Agent, source IP, ASN, reverse DNS, method, path, status, response size, redirects, timestamp, and request rate. The User-Agent often includes the specific App Engine Application ID (appid: [APP_ID]), which can help you identify if multiple requests are coming from the same third-party app.

Traffic will originate from Google Cloud IP ranges. However, verifying that the IP belongs to Google Cloud only confirms the hosting provider, not the legitimacy of the request.

Analyze behavior to determine intent. High concurrency, repeated retries, deep traversal of non-public areas, or extraction of proprietary data indicates scraping. Since you cannot rely on robots.txt compliance, you must use active enforcement.

Evaluate /robots.txt independently. Confirm the canonical host, response status, content type, exact user-agent group, and path match. A page-level directive may express a discoverability preference:

configuration / code
<meta name="robots" content="noindex, nofollow">
configuration / code
X-Robots-Tag: noindex, nofollow

These signals do not establish an opt-out for an undocumented client and do not secure private routes. Use authenticated delivery, signed URLs, and application authorization.

WAF and Nginx remediation examples

Because AppEngine-Google traffic is often unverified scraping, blocking it at the edge is a common and effective remediation if you do not expect legitimate webhooks or API calls from App Engine hosted services.

A narrow WAF rule can block the declared identity. Replace the example expression if your observed header differs:

configuration / code
{
  "description": "Block AppEngine-Google third-party traffic",
  "expression": "lower(http.user_agent) contains \"appengine-google\"",
  "action": "block"
}

For Nginx, you can block this User-Agent globally or scope enforcement to specific routes:

configuration / code
map $http_user_agent $block_appengine_google {
    default 0;
    ~*AppEngine-Google 1;
}

server {
    # Block globally for the entire server
    if ($block_appengine_google) { return 403; }

    location / {
        try_files $uri $uri/ =404;
    }
}

A User-Agent rule is easy to spoof or evade. If the scraper simply changes their User-Agent, you will need to rely on rate limiting, anomaly detection, and potentially blocking specific Google Cloud IP ranges if the abuse persists. Pair edge matching with authentication and signed assets.

Review checklist

Search logs for every exact header that may be associated with AppEngine-Google and preserve representative requests. Note the appid if present to identify specific aggressive actors. Record paths, response sizes, statuses, source networks, timing, and rate.

Decide whether your objective is to allow specific webhooks, prevent extraction, protect private content, or reduce crawl load. Since robots.txt is largely ineffective for this User-Agent, enforce sensitive routes with WAF and application controls, and test docs, media, feeds, sitemaps, uploads, and APIs separately.

References

  1. Google App Engine Documentation — general documentation for the hosting platform generating this traffic.
  2. Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
  3. RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.