Bot directory / scraper

adidxbot: Robots.txt & Crawl Policy Reference

Technical reference for adidxbot, Microsoft's advertising index crawler. Learn how it interacts with landing pages and how to manage its access.

AI Summary: adidxbot is a web crawler operated by Microsoft for Bing Ads (Microsoft Advertising). It visits ad landing pages to evaluate their quality, relevance, and compliance with advertising policies. Blocking this bot may result in your ads being disapproved or paused. Treat it as authorized traffic if you are running Microsoft Advertising campaigns.

Role and policy boundary

The registry describes adidxbot as Bing's advertising index web crawler. While the specific Bing Webmaster Tools documentation page was inaccessible during headless-browser review, adidxbot is a well-known crawler used by Microsoft Advertising. Its primary purpose is to follow the URLs specified in your ads to verify that the landing pages are active, relevant, and policy-compliant.

Because its purpose is ad verification rather than general search indexing or AI training, blocking it can directly impact your active campaigns. Do not infer that it collects AI-training data or builds the general Bing search index (which is handled by Bingbot).

If your logs confirm an exact adidxbot token and you want to communicate a restriction (though generally not recommended for active ad landing pages), publish:

configuration / code
User-agent: adidxbot
Disallow: /

For selective access (e.g., allowing it only on landing pages):

configuration / code
User-agent: adidxbot
Allow: /landing-pages/
Disallow: /private/
Disallow: /internal/
Disallow: /api/

Robots.txt is advisory and cannot protect private or licensed content. Use authentication, authorization, signed URLs, and origin controls for those boundaries.

Layered verification

Start with raw access logs and preserve the complete User-Agent, source IP, ASN, reverse DNS, method, path, status, response size, redirects, timestamp, and request rate. Microsoft typically provides ways to verify their crawlers via reverse DNS (e.g., checking for search.msn.com or similar Microsoft domains), though you should consult current Microsoft documentation for the exact verification steps for adidxbot.

Analyze behavior without assigning purpose prematurely. Requests for ad landing pages resemble authorized verification; deep traversal of non-ad areas, high concurrency, repeated retries, original-asset downloads, or private API access may indicate misconfiguration or spoofing. These patterns demonstrate operational impact but cannot prove the operator or downstream use.

Evaluate /robots.txt independently. Confirm the canonical host, response status, content type, exact user-agent group, and path match. A page-level directive may express a discoverability preference:

configuration / code
<meta name="robots" content="noindex, nofollow">
configuration / code
X-Robots-Tag: noindex, nofollow

These signals do not establish an opt-out for an undocumented client and do not secure private routes. Use authenticated delivery, signed URLs, and application authorization.

WAF and Nginx remediation examples

Once logs confirm an exact unwanted token (e.g., a spoofed bot or if you are not running ads), a narrow WAF rule can block the declared identity. Replace the example expression if your observed header differs:

configuration / code
{
  "description": "Block observed adidxbot token",
  "expression": "lower(http.user_agent) contains \"adidxbot\"",
  "action": "block"
}

For Nginx, scope enforcement to private and high-cost routes while investigating public access:

configuration / code
map $http_user_agent $block_adidxbot {
    default 0;
    ~*adidxbot 1;
}

server {
    location ~ ^/(private|internal|account|uploads|paywall|api)/ {
        if ($block_adidxbot) { return 403; }
        try_files $uri $uri/ =404;
    }
}

A User-Agent rule is easy to spoof or evade and may block a legitimate client using the substring. Do not create an IP allowlist or broad network block without current operator evidence. Test browsers, social previews, feed readers, search crawlers, approved monitors, and customer integrations. Pair edge matching with authentication, rate limits, signed assets, and anomaly detection.

Review checklist

Search logs for every exact header that may be associated with adidxbot and preserve representative requests. Record paths, response sizes, statuses, source networks, timing, and rate. Verify the traffic against Microsoft's published IP ranges or reverse DNS methods if available.

Decide whether your objective is to preserve ad campaign functionality, prevent extraction, protect private content, or reduce crawl load. Publish a targeted robots group only for the exact observed token, enforce sensitive routes with WAF and application controls, and test docs, media, feeds, sitemaps, uploads, and APIs separately. Revisit the profile if Microsoft updates the specific documentation for this crawler.

References

  1. Registry-linked Bing Webmaster Tools policy URL — the content was inaccessible during the 2026-08-24 headless-browser review.
  2. Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
  3. RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.