Bot directory / search-engine

Qwantify: Robots.txt & Crawl Policy Reference

Technical reference for the Qwantify registry label, separating it from Qwant’s currently documented Qwantbot family and verification process.

AI Summary: The registry labels Qwantify as a Qwant search crawler and records Mozilla/5.0 (compatible; Qwantify/2.0n; +https://www.qwant.com/)/*, but Qwant’s current first-party crawler page documents Qwantbot and Qwantbot-news, not Qwantify. Qwant’s product homepage is current, yet no Qwantify-specific policy, IP range, DNS verification, rate, or opt-out was established. Treat Qwantify as an alternate or historical registry label and verify any observed request independently.

Role and policy boundary

The registry describes Qwantify as a Qwant search-engine web crawler and records this User-Agent:

configuration / code
Mozilla/5.0 (compatible; Qwantify/2.0n; +https://www.qwant.com/)/*

The current Qwant homepage confirms the search product is active and links to its Help Center. Qwant’s first-party crawler page, reviewed for the related Qwantbot profile, documents the Qwantbot and Qwantbot-news families and their verification method. It does not document Qwantify as a current crawler token. Do not silently merge the registry label with the documented Qwantbot family or “correct” the registry header into a different value.

The registry value also ends with /*, which is not evidence that a client sends that literal suffix. Preserve the complete observed header and compare it with the registry candidate. A matching request could be a legacy deployment, test client, browser-like tool, unrelated Qwant reference, or spoofed header. The token does not establish search indexing, AI input, model training, data retention, or permission to access private content.

Because no Qwantify-specific policy was verified, do not publish a Qwantify robots group as if it came from Qwant. If logs establish a current, attributable token and you decide to exclude it, use the exact observed value in a deliberate site-owner rule:

configuration / code
User-agent: CONFIRMED-QWANTIFY
Disallow: /

For selective public access after confirmation:

configuration / code
User-agent: CONFIRMED-QWANTIFY
Allow: /public/
Allow: /docs/
Disallow: /admin/
Disallow: /account/
Disallow: /private/
Disallow: /licensed/
Disallow: /api/

These are explanatory site-owner controls, not recovered Qwantify instructions. Robots.txt is advisory and cannot secure private or licensed content; use authentication, authorization, signed URLs, data minimization, and origin controls.

Layered verification

Start with raw access logs and preserve the complete User-Agent, source IP, ASN, reverse-DNS result, HTTP method, requested path, response status, response size, redirect chain, timestamp, concurrency, and request rate. No Qwantify-specific first-party source in this review published an IP range, reverse-DNS procedure, Crawl-delay, rate guidance, removal address, or verification workflow.

Qwant’s documented Qwantbot process is not automatically a Qwantify process. If the client claims Qwant ownership, use the current Qwant crawler documentation and contact path to ask whether the alternate token is recognized. The documented method for Qwantbot—reverse DNS ending with qwant.com, optional forward confirmation, and a refreshable JSON prefix—must not be represented as Qwantify verification without operator confirmation.

Do not treat the qwant.com URL in the registry header as proof of ownership. Check source IP and DNS evidence independently. A User-Agent match, a Qwant-associated IP, or a reverse-DNS suffix without forward confirmation is insufficient to authorize a request.

Compare observed behavior with a search-crawling hypothesis without turning it into attribution. Public HTML, metadata, feeds, sitemaps, and ordinary assets may be consistent with indexing. Private endpoints, licensed content, account routes, APIs, high concurrency, repeated retries, or unexpected bulk downloads establish impact and load risk, not Qwantify identity or downstream use.

Evaluate /robots.txt only after the actual observed token is known. Confirm that your host returns a successful text response and contains the exact group you intend to publish. The current Qwant crawler page says Qwantbot crawlers respect the robots standard, but that statement does not authenticate Qwantify. Page-level directives can express indexing preferences:

configuration / code
<meta name="robots" content="noindex, nofollow">
configuration / code
X-Robots-Tag: noindex, nofollow

These signals do not authenticate an undocumented crawler or secure private paths. Enforce sensitive boundaries in the application and at the origin. If your policy distinguishes search indexing, AI input, reference use, and model training, document each purpose separately.

WAF and Nginx remediation examples

When the identity is unverified, use report-only logging and a narrow observation. Avoid broad Qwant, Qwantbot, browser, or search rules that can affect unrelated clients:

configuration / code
{
  "description": "Observe unverified Qwantify candidates",
  "expression": "lower(http.user_agent) contains \"qwantify\"",
  "action": "log"
}

After independent verification and a policy decision, scope enforcement to sensitive routes and preserve evidence supporting the rule:

configuration / code
map $http_user_agent $block_confirmed_qwantify_private {
    default 0;
    # Add only a complete, independently verified token here.
    # ~*Qwantify/2\.0n 1;
}

server {
    location ~ ^/(admin|account|private|licensed|internal|api)/ {
        if ($block_confirmed_qwantify_private) { return 403; }
        try_files $uri $uri/ =404;
    }
}

A User-Agent match is easy to spoof and the registry token is not a documented current Qwant family value. Do not invent an IP allowlist, Qwant ASN exception, reverse-DNS suffix, rate, training policy, or permanent trust rule. Use per-client rate limits, concurrency ceilings, timeouts, response-size controls, caching, and anomaly detection at the edge or origin.

Start in report-only mode, review false positives, and restrict only after evidence supports the action. Test public pages, sitemaps, structured data, licensed assets, account routes, APIs, 429 behavior, response-size limits, and approved integrations separately. Pair WAF controls with authentication and application authorization.

Review checklist

Search logs for the exact Qwantify substring and preserve the complete header, source IP, ASN, PTR result, forward lookup, path, method, response size, status, timing, rate, concurrency, and redirects. Compare it with the registry candidate without claiming it is a current Qwantbot.

Re-check Qwant’s current crawler page, Qwant homepage, robots documentation, IP JSON, and contact path. The reviewed first-party crawler source documents Qwantbot/Qwantbot-news rather than Qwantify. Keep this profile at documented-limit until Qwant publishes or confirms the exact alternate token, purpose, robots behavior, source-verification method, rate guidance, or opt-out process.

Check your robots file for an exact Qwantify group only if the token is confirmed. Do not apply Qwantbot’s documented IP/DNS method to Qwantify without operator confirmation. Use META or X-Robots-Tag for indexing preferences and authentication for private resources.

Review search indexing, AI input, reference use, and model-training decisions separately. Qwant’s search product and Qwantbot documentation do not establish downstream permission for an alternate registry token. Do not claim successful blocking or verification from configuration alone; validate later logs and obtain a current operator response.

Decide whether your objective is to preserve public discovery, reduce crawl load, limit extraction, protect licensed material, or prevent private access. Publish exact robots rules only after the token is known and enforce sensitive routes at the application and origin layers.

References

  1. Qwant — current first-party product homepage, reviewed with the headless browser on 2026-08-25; it confirms the search service but does not document Qwantify.
  2. Qwant Web crawler — first-party crawler page documenting Qwantbot/Qwantbot-news, robots behavior, DNS verification, IP JSON, and contact; it was not used to upgrade Qwantify to a current official token.
  3. Qwantbot IP JSON — current Qwantbot-family snapshot; not treated as a Qwantify allowlist.
  4. Crawler User Agents community registry — registry context for Qwantify; it does not authenticate current Qwant infrastructure.
  5. Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
  6. RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.