Bot directory / search-engine

Gigabot: Robots.txt & Crawl Policy Reference

Technical reference for the historical Gigabot label associated with Gigablast, with explicit limits around its unavailable documentation and current verification.

AI Summary: Gigabot/1.0 is a historical registry label associated with Gigablast. The registry documentation URL did not resolve over either HTTP or HTTPS during review, and no current first-party policy, IP range, or verification method was available. Treat matching traffic as an unverified User-Agent observation, not proof of an active official Gigablast crawler.

Role and policy boundary

The registry describes Gigabot as the web crawler for the Gigablast search engine and records the User-Agent Gigabot/1.0. That label supplies a useful log-search term, but it does not prove current operation. The linked documentation URL, http://www.gigablast.com/spider.html, failed with net::ERR_NAME_NOT_RESOLVED in the headless-browser review; the HTTPS equivalent failed the same way. No current first-party page was available to confirm an operator, production service, robots behavior, rate policy, or source network.

Gigablast also maintains a public open-source search-engine repository, but source code and repository availability cannot authenticate a request from a deployed crawler. A fork, research instance, independent operator, or spoofed header may use the same token. Do not infer current activity, downstream use, AI-training purpose, or permission to crawl from the name alone.

If your logs show this exact token and you have decided to exclude it from search discovery, a narrow site-owner rule could be:

configuration / code
User-agent: Gigabot/1.0
Disallow: /

For a selective policy, make the permitted and restricted paths explicit:

configuration / code
User-agent: Gigabot/1.0
Allow: /public/
Allow: /docs/
Disallow: /admin/
Disallow: /account/
Disallow: /private/
Disallow: /api/

These are example controls, not recovered Gigablast instructions. Robots.txt is advisory and cannot protect private or licensed content. Use authentication, authorization, signed URLs, and origin controls for those boundaries.

Layered verification

Begin with access logs and capture the complete User-Agent, source IP, ASN, reverse-DNS result, HTTP method, requested path, response status, response size, redirect chain, timestamp, and request rate. Since the reviewed documentation domain did not resolve and published no verification method, the header alone must remain untrusted input.

Compare observed behavior with a search-spider hypothesis without treating it as attribution. Public HTML, canonical metadata, feeds, sitemaps, and ordinary assets may be consistent with indexing. Requests to private endpoints, aggressive concurrency, repeated retries, unusual downloads, or traffic that disregards your published policy may indicate spoofing, abuse, a fork, or an unrelated client. These observations establish operational impact, not identity or downstream use.

Evaluate /robots.txt independently. Confirm that it is served by the intended host, returns a successful status and text content type, and contains the exact group you intended to publish. Test whether the parser handles the full token and whether a global rule changes the outcome. Page-level directives can express discovery preferences:

configuration / code
<meta name="robots" content="noindex, nofollow">
configuration / code
X-Robots-Tag: noindex, nofollow

Neither directive authenticates Gigabot or protects a private route. Enforce sensitive boundaries in the application and at the origin, and do not claim that an unavailable operator has accepted your opt-out.

WAF and Nginx remediation examples

If logs show a repeatable unwanted token, start with a report-only rule. Adapt the expression to your provider and label the event as an observed header, not a verified official crawler:

configuration / code
{
  "description": "Review observed Gigabot traffic",
  "expression": "lower(http.user_agent) contains \"gigabot\"",
  "action": "log"
}

After reviewing false positives and confirming the business decision, scope enforcement to sensitive routes:

configuration / code
map $http_user_agent $review_gigabot {
    default 0;
    ~*Gigabot/1\.0 1;
}

server {
    location ~ ^/(admin|account|private|internal|api)/ {
        if ($review_gigabot) { return 403; }
        try_files $uri $uri/ =404;
    }
}

Do not invent an IP allowlist or reverse-DNS suffix for Gigabot. A User-Agent is easy to spoof and a broad case-insensitive match can catch an authorized tool or a different historical client. Test public pages, feeds, sitemaps, media, uploads, account flows, APIs, and approved integrations separately. Pair edge controls with authentication, rate limits, signed assets, caching, and anomaly detection.

Review checklist

Search logs for the exact Gigabot/1.0 token and preserve representative requests, source IPs, ASNs, reverse-DNS results, paths, methods, response sizes, statuses, timing, and rate. Check whether the traffic is reproducible and whether any independently verified infrastructure is associated with it. Do not classify a request as authentic from its header, repository name, or historical documentation URL.

Re-check the registry link and the related Gigablast repository for a new canonical operator source. Keep this profile at legacy-label until an operator publishes a current policy, canonical User-Agent, source-verification method, or robots guidance. Decide whether the objective is to preserve search visibility, limit extraction, protect private material, or reduce load, and choose exact robots, WAF, and application controls accordingly.

Do not claim successful blocking from configuration alone. Verify subsequent access logs, response statuses, cache behavior, and legitimate search impact. If a future source identifies a different active crawler token, record it as new evidence rather than silently treating it as Gigabot/1.0.

References

  1. Gigablast spider documentation URL — direct HTTP headless-browser request failed with net::ERR_NAME_NOT_RESOLVED on 2026-08-25.
  2. Gigablast spider documentation HTTPS equivalent — direct HTTPS headless-browser request failed with net::ERR_NAME_NOT_RESOLVED on 2026-08-25.
  3. Gigablast open-source search engine repository — related public code project; it does not authenticate current Gigabot traffic.
  4. Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
  5. RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.