Bot directory / ai-training

www.spider: Robots.txt & Crawl Policy Reference

Technical reference for the www.spider registry label associated with Spider.com's commercial real-time web-data extraction service. Learn how to verify traffic without inferring a fixed crawler identity.

AI Summary: Spider.com's official page documents a commercial Real-Time Crawler and proxy-backed web-data extraction service, but it does not publish a dedicated Spider crawler contract, source ranges, or robots behavior. The registry User-Agent is therefore only a product-associated observation. Verify actual traffic from logs and protect sensitive routes with authentication, rate limits, and narrow edge controls.

Role and policy boundary

Spider.com describes itself as a premium proxy provider specializing in automated web-data extraction. Its Real-Time Crawler is advertised as customizable, able to handle CAPTCHAs, and supported by residential proxy infrastructure and IP rotation. The page lists crawl-and-index, market research, SEO monitoring, brand protection, travel aggregation, and geo-blocked data among its use cases.

This establishes a commercial extraction product, not a fixed vendor crawler operating under one header against every website. The registry associates Mozilla/5.0 (compatible; Spider; +https://www.spider.com/) with Spider.com, but the official page reviewed does not publish that exact User-Agent, source IP ranges, crawl rate, robots behavior, or a webmaster verification method. A customer using Spider's service may also generate requests through rotating residential addresses, so a browser-like header and vendor domain are not sufficient identity proof.

Do not infer AI-training use from the registry category or from the existence of an extraction API. If your logs confirm the exact token and you want to communicate a restriction, publish:

configuration / code
User-agent: Spider
Disallow: /

For selective access:

configuration / code
User-agent: Spider
Allow: /public-reference/
Allow: /docs/
Disallow: /private/
Disallow: /originals/
Disallow: /uploads/
Disallow: /api/

Robots.txt is advisory. It cannot protect private or licensed content from a proxy-backed requester; use authentication, authorization, signed URLs, and origin controls.

Layered verification

Start with raw access logs and preserve the complete User-Agent, source IP, ASN, reverse DNS, method, path, status, response size, redirects, timestamp, and request rate. The official product page says Spider offers residential proxies and overcoming IP blocking, which makes source-IP attribution particularly difficult. Do not treat a residential IP or the Spider token as conclusive identity evidence.

Analyze behavior without assigning purpose prematurely. One-off URL extraction, deep traversal, media retrieval, high concurrency, repeated retries, CAPTCHA-related patterns, and requests for private APIs have different operational implications. They can demonstrate load or access-control risk, but cannot prove that Spider.com itself made the request or that the data is being used for AI training.

Evaluate /robots.txt independently. Confirm the canonical host, status, content type, exact group, and path match. Because the official page does not establish Spider's robots behavior, an allowed request is not evidence of compliance. Page-level metadata may express a discoverability preference:

configuration / code
<meta name="robots" content="noindex, nofollow">
configuration / code
X-Robots-Tag: noindex, nofollow

These signals do not secure private routes and may not be honored by a configurable extraction client. Enforce sensitive boundaries in the application and at the origin.

WAF and Nginx remediation examples

If logs confirm unwanted requests carrying the token, use an exact, narrow WAF rule and monitor false positives:

configuration / code
{
  "description": "Block observed Spider crawler token",
  "expression": "lower(http.user_agent) contains \"compatible; spider;\"",
  "action": "block"
}

For Nginx, scope enforcement to private, original-asset, and high-cost routes while investigating public access:

configuration / code
map $http_user_agent $block_spider_token {
    default 0;
    ~*compatible;[[:space:]]+spider; 1;
}

server {
    location ~ ^/(private|originals|uploads|internal|paywall|api)/ {
        if ($block_spider_token) { return 403; }
        try_files $uri $uri/ =404;
    }
}

A User-Agent rule is easy to spoof or evade and may block an approved integration. Do not create a broad IP denylist from residential proxy traffic or a third-party list. Test browsers, social previews, feed readers, search crawlers, approved monitors, and customer integrations. Pair edge matching with route authorization, rate limits, CAPTCHA or challenge controls where appropriate, and signed asset delivery.

Review checklist

Search logs for the exact Spider declaration and preserve source network, path, method, response size, status, timing, and concurrency. Compare the observed pattern with Spider.com's current Real-Time Crawler description, but keep the product association separate from proof of operator identity. During this review no dedicated Spider User-Agent or source-verification contract was published on the official crawler page.

Decide whether your objective is to preserve public discovery, prevent extraction of licensed media, reduce bandwidth, or secure private APIs. Publish a targeted robots group for the exact observed token, enforce private and high-cost routes with WAF and application authorization, and test pages, media, feeds, sitemaps, uploads, and APIs separately. Revisit the profile if Spider.com publishes a crawler contract, source verification, robots policy, or customer-controlled User-Agent guidance.

References

  1. Spider Real-Time Crawler — official product page documenting automated web-data extraction, proxy infrastructure, customization, and use cases; it does not publish a dedicated crawler contract.
  2. Spider API Documentation — official documentation link exposed on Spider.com; API access should not be conflated with a fixed public crawler identity.
  3. Google Robots.txt Introduction — general explanation of crawler directives and their limitations.

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.