Bot directory / search-engine

GigablastOpenSource: Robots.txt & Crawl Policy Reference

Technical reference for the Gigablast open-source search spider label, with evidence limits around the current User-Agent, robots behavior, and production verification.

AI Summary: Gigablast’s public repository documents an open-source distributed search engine and spider/crawler written in C/C++ for Linux. It does not verify that GigablastOpenSource/1.0 is a current production User-Agent, and the historical gigablast.com host did not resolve during review. Treat this as a partially documented software and traffic label: identify requests from logs before adding robots or WAF controls.

Role and policy boundary

The linked first-party GitHub repository describes Gigablast as a distributed open-source search engine and spider/crawler written in C/C++ for Linux on Intel/AMD. That establishes the project’s technical role, but it is not proof that a request bearing GigablastOpenSource/1.0 comes from an operator-controlled deployment. The repository title references November 2017, and its latest visible commit at review was an Update LICENSE change dated January 10, 2024. A public code repository can be used by independent operators, forks, research deployments, or unrelated clients.

The registry User-Agent is GigablastOpenSource/1.0. No current first-party page reviewed in this pass confirmed that exact header, published a production bot name, listed crawler IP ranges, or explained an operator-managed opt-out. The direct request to https://www.gigablast.com/robots.txt failed with net::ERR_NAME_NOT_RESOLVED, so no current site policy was inferred from it. Do not call a request authentic solely because it includes the registry token.

If access logs confirm that this exact token is used by a deployment you want excluded from discovery, a narrow robots group could be written as:

configuration / code
User-agent: GigablastOpenSource/1.0
Disallow: /

For a selective public-search boundary, replace the paths with the routes you have deliberately approved:

configuration / code
User-agent: GigablastOpenSource/1.0
Allow: /public/
Allow: /docs/
Disallow: /admin/
Disallow: /account/
Disallow: /private/
Disallow: /api/

This configuration is a site-owner choice, not a published Gigablast policy. Robots.txt is advisory and cannot protect private or licensed material; use authentication, authorization, signed URLs, and origin controls for those boundaries. Do not assume that source code, repository ownership, or a User-Agent creates permission to crawl.

Layered verification

Start with access logs and preserve the complete User-Agent, source IP, ASN, reverse-DNS result, HTTP method, requested path, response status, response size, redirect chain, timestamp, and request rate. Because the reviewed source did not publish a current crawler network or DNS verification method, treat the observed header as a lead rather than an identity assertion.

Compare request behavior with a search-spider hypothesis without presenting that hypothesis as fact. Requests for public HTML, canonical metadata, feeds, sitemaps, and ordinary page assets may be consistent with indexing. Private endpoint access, high-concurrency traversal, repeated retries, large unexpected downloads, or traffic that ignores your published restrictions may indicate spoofing, a fork, abuse, or a different client. These observations establish impact and operational risk, not the operator’s identity or downstream use.

Evaluate /robots.txt independently. Confirm the response is served from the correct host, returns a successful status and text content type, and contains the exact group you intend to use. Test matching behavior for the full token, a shorter token only if intentionally chosen, and the global group. Page-level directives can express discovery preferences:

configuration / code
<meta name="robots" content="noindex, nofollow">
configuration / code
X-Robots-Tag: noindex, nofollow

These signals do not authenticate the crawler, secure a private route, or prove that an open-source deployment has accepted your preference. Enforce sensitive boundaries in the application and at the origin.

WAF and Nginx remediation examples

If logs show a repeatable and unwanted request token, begin with a report-only rule. Adapt the expression to your provider’s syntax and do not claim that it identifies an official Gigablast service:

configuration / code
{
  "description": "Review observed GigablastOpenSource traffic",
  "expression": "lower(http.user_agent) contains \"gigablastopensource\"",
  "action": "log"
}

After reviewing false positives and confirming the business decision, an edge rule may restrict sensitive routes:

configuration / code
map $http_user_agent $review_gigablast_open_source {
    default 0;
    ~*GigablastOpenSource 1;
}

server {
    location ~ ^/(admin|account|private|internal|api)/ {
        if ($review_gigablast_open_source) { return 403; }
        try_files $uri $uri/ =404;
    }
}

A User-Agent match is easy to spoof and can also catch an authorized internal test or a third-party fork. Do not create an IP allowlist from the repository or from DNS guesses. Test public pages, feeds, sitemaps, media, uploads, account flows, APIs, and approved integrations separately. Pair edge rules with authentication, rate limits, signed assets, caching, and anomaly detection.

Review checklist

Search logs for the complete GigablastOpenSource/1.0 token and record representative requests, source IPs, ASNs, reverse-DNS results, paths, methods, response sizes, statuses, timing, and rate. Check whether the traffic is reproducible and whether it is coming from an infrastructure source that you can verify independently. Do not mark a request authentic merely because it matches the registry value.

Re-check the official repository for updated documentation and retry the historical website only when you have a confirmed replacement domain. Keep the profile at partially-documented until an operator publishes a current crawler policy, canonical User-Agent, source-verification method, or robots guidance. Decide whether you want to preserve search visibility, limit extraction, protect private material, or reduce load; then publish exact robots rules and enforce private routes with application and WAF controls.

Do not claim successful blocking from a configuration change alone. Verify subsequent access logs, response statuses, cache behavior, and the effect on legitimate search visibility. If a new deployment uses a different header, create or update the corresponding evidence rather than silently broadening this profile.

References

  1. Gigablast open-source search engine repository — first-party public repository describing the distributed search engine and spider/crawler, reviewed 2026-08-25.
  2. Gigablast historical robots URL — direct headless-browser request failed with net::ERR_NAME_NOT_RESOLVED during review; no policy was inferred.
  3. Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
  4. RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.