GigablastOpenSource: Robots.txt & Crawl Policy Reference
Technical reference for the Gigablast open-source search spider label, with evidence limits around the current User-Agent, robots behavior, and production verification.
AI Summary: Gigablast’s public repository documents an open-source distributed search engine and spider/crawler written in C/C++ for Linux. It does not verify that
GigablastOpenSource/1.0is a current production User-Agent, and the historicalgigablast.comhost did not resolve during review. Treat this as a partially documented software and traffic label: identify requests from logs before adding robots or WAF controls.
Role and policy boundary
The linked first-party GitHub repository describes Gigablast as a distributed open-source search engine and spider/crawler written in C/C++ for Linux on Intel/AMD. That establishes the project’s technical role, but it is not proof that a request bearing GigablastOpenSource/1.0 comes from an operator-controlled deployment. The repository title references November 2017, and its latest visible commit at review was an Update LICENSE change dated January 10, 2024. A public code repository can be used by independent operators, forks, research deployments, or unrelated clients.
The registry User-Agent is GigablastOpenSource/1.0. No current first-party page reviewed in this pass confirmed that exact header, published a production bot name, listed crawler IP ranges, or explained an operator-managed opt-out. The direct request to https://www.gigablast.com/robots.txt failed with net::ERR_NAME_NOT_RESOLVED, so no current site policy was inferred from it. Do not call a request authentic solely because it includes the registry token.
If access logs confirm that this exact token is used by a deployment you want excluded from discovery, a narrow robots group could be written as:
User-agent: GigablastOpenSource/1.0
Disallow: /
For a selective public-search boundary, replace the paths with the routes you have deliberately approved:
User-agent: GigablastOpenSource/1.0
Allow: /public/
Allow: /docs/
Disallow: /admin/
Disallow: /account/
Disallow: /private/
Disallow: /api/
This configuration is a site-owner choice, not a published Gigablast policy. Robots.txt is advisory and cannot protect private or licensed material; use authentication, authorization, signed URLs, and origin controls for those boundaries. Do not assume that source code, repository ownership, or a User-Agent creates permission to crawl.
Layered verification
Start with access logs and preserve the complete User-Agent, source IP, ASN, reverse-DNS result, HTTP method, requested path, response status, response size, redirect chain, timestamp, and request rate. Because the reviewed source did not publish a current crawler network or DNS verification method, treat the observed header as a lead rather than an identity assertion.
Compare request behavior with a search-spider hypothesis without presenting that hypothesis as fact. Requests for public HTML, canonical metadata, feeds, sitemaps, and ordinary page assets may be consistent with indexing. Private endpoint access, high-concurrency traversal, repeated retries, large unexpected downloads, or traffic that ignores your published restrictions may indicate spoofing, a fork, abuse, or a different client. These observations establish impact and operational risk, not the operator’s identity or downstream use.
Evaluate /robots.txt independently. Confirm the response is served from the correct host, returns a successful status and text content type, and contains the exact group you intend to use. Test matching behavior for the full token, a shorter token only if intentionally chosen, and the global group. Page-level directives can express discovery preferences:
<meta name="robots" content="noindex, nofollow">
X-Robots-Tag: noindex, nofollow
These signals do not authenticate the crawler, secure a private route, or prove that an open-source deployment has accepted your preference. Enforce sensitive boundaries in the application and at the origin.
WAF and Nginx remediation examples
If logs show a repeatable and unwanted request token, begin with a report-only rule. Adapt the expression to your provider’s syntax and do not claim that it identifies an official Gigablast service:
{
"description": "Review observed GigablastOpenSource traffic",
"expression": "lower(http.user_agent) contains \"gigablastopensource\"",
"action": "log"
}
After reviewing false positives and confirming the business decision, an edge rule may restrict sensitive routes:
map $http_user_agent $review_gigablast_open_source {
default 0;
~*GigablastOpenSource 1;
}
server {
location ~ ^/(admin|account|private|internal|api)/ {
if ($review_gigablast_open_source) { return 403; }
try_files $uri $uri/ =404;
}
}
A User-Agent match is easy to spoof and can also catch an authorized internal test or a third-party fork. Do not create an IP allowlist from the repository or from DNS guesses. Test public pages, feeds, sitemaps, media, uploads, account flows, APIs, and approved integrations separately. Pair edge rules with authentication, rate limits, signed assets, caching, and anomaly detection.
Review checklist
Search logs for the complete GigablastOpenSource/1.0 token and record representative requests, source IPs, ASNs, reverse-DNS results, paths, methods, response sizes, statuses, timing, and rate. Check whether the traffic is reproducible and whether it is coming from an infrastructure source that you can verify independently. Do not mark a request authentic merely because it matches the registry value.
Re-check the official repository for updated documentation and retry the historical website only when you have a confirmed replacement domain. Keep the profile at partially-documented until an operator publishes a current crawler policy, canonical User-Agent, source-verification method, or robots guidance. Decide whether you want to preserve search visibility, limit extraction, protect private material, or reduce load; then publish exact robots rules and enforce private routes with application and WAF controls.
Do not claim successful blocking from a configuration change alone. Verify subsequent access logs, response statuses, cache behavior, and the effect on legitimate search visibility. If a new deployment uses a different header, create or update the corresponding evidence rather than silently broadening this profile.
References
- Gigablast open-source search engine repository — first-party public repository describing the distributed search engine and spider/crawler, reviewed 2026-08-25.
- Gigablast historical robots URL — direct headless-browser request failed with
net::ERR_NAME_NOT_RESOLVEDduring review; no policy was inferred. - Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
- RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.
Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.