← Bot Directory/Linguee Bot
Bot directory / search-engine

Linguee Bot: Robots.txt & Crawl Policy Reference

Technical reference for the historical Linguee Bot label, including the current robots.txt training restriction and limits around its unavailable bot page.

AI Summary: The registry associates Linguee Bot with translation-search crawling, but its linked bot page currently returns 404 Not Found. Linguee’s official robots.txt contains a prominent restriction against using crawled data to train machine-translation systems and extensive product-specific rules, yet the reviewed file did not publish a dedicated Linguee Bot group, IP range, or source-verification procedure. Treat the registry header as unverified and preserve the training restriction as a site policy signal, not authentication.

Role and policy boundary

The inventory records this historical User-Agent:

configuration / code
Linguee Bot (http://www.linguee.com/bot)

The registry describes it as a translation web crawler, but the linked http://www.linguee.com/bot page redirected to HTTPS and returned 404 Not Found in the headless-browser review. That means the current bot identity, product role, rate policy, and operator instructions were not confirmed by a dedicated page.

The current official https://www.linguee.com/robots.txt begins with an explicit warning that crawled Linguee data must not be used to train machine-translation systems. It also warns that Linguee contains fake entries that can help identify copying without permission. This is a strong site policy signal for downstream use and content protection, but it does not authenticate a requester or prove that every client will obey it.

Do not infer current Linguee Bot activity, operator identity, search indexing, AI training, or permission to access private translation content from the historical header. A matching request may be a legacy client, a partner, a test harness, or a spoofed token.

If your logs confirm the exact token and your site policy is to exclude it from discovery, a narrow defensive group could be:

configuration / code
User-agent: Linguee Bot
Disallow: /

For selective access to approved public translation pages while protecting licensed or sensitive content:

configuration / code
User-agent: Linguee Bot
Allow: /public/
Allow: /docs/
Disallow: /private/
Disallow: /licensed/
Disallow: /account/
Disallow: /api/

These are site-owner examples, not recovered Linguee instructions. Robots.txt is advisory and cannot protect private or licensed content; use authentication, authorization, signed URLs, and origin controls for those boundaries. If you publish a training opt-out, make it explicit in policy and retain evidence of the exact version and date.

Layered verification

Begin with raw access logs and preserve the complete User-Agent, source IP, ASN, reverse-DNS result, HTTP method, requested path, response status, response size, redirect chain, timestamp, and request rate. The reviewed robots file publishes extensive rules but no dedicated Linguee Bot IP range, reverse-DNS verification procedure, or Crawl-delay. The historical header is therefore a triage signal, not authentication.

Compare observed behavior with a translation-search hypothesis without turning it into attribution. Requests for public bilingual pages, phrase pages, sitemaps, and ordinary assets may be consistent with indexing. Requests for licensed dictionaries, private translation projects, accounts, APIs, or unusual bulk downloads may indicate an authorization problem, scraping, spoofing, or another client. These observations establish impact and data risk, not operator identity or downstream use.

Evaluate /robots.txt independently. Confirm that it is served from the correct host, returns a successful status and text content type, and contains the exact group you intend to publish. The current Linguee file includes separate groups for Googlebot, Mediapartners-Google, Yahoo! Slurp, and other named agents; do not assume those groups govern the historical Linguee Bot token. Test exact matching under your parser and keep the training restriction separate from crawl access.

Page-level directives can express discovery preferences:

configuration / code
<meta name="robots" content="noindex, nofollow">
configuration / code
X-Robots-Tag: noindex, nofollow

For images and licensed text, use the appropriate resource-level policy and contractual controls in addition to these signals. They do not authenticate a bot or secure private routes. Enforce sensitive boundaries in the application and at the origin.

WAF and Nginx remediation examples

If logs show the historical token, begin with a report-only rule and correlate it with source evidence and the data being requested. Adapt the expression to your WAF provider:

configuration / code
{
  "description": "Review Linguee Bot and translation-content requests",
  "expression": "lower(http.user_agent) contains \"linguee bot\"",
  "action": "log"
}

After reviewing false positives and confirming the business decision, scope enforcement to licensed or private routes:

configuration / code
map $http_user_agent $block_linguee_private {
    default 0;
    ~*Linguee[ ]Bot 1;
}

server {
    location ~ ^/(private|licensed|account|internal|api)/ {
        if ($block_linguee_private) { return 403; }
        try_files $uri $uri/ =404;
    }
}

A User-Agent match is easy to spoof and should not be used as the sole control for high-value translation data. Do not invent an IP allowlist or reverse-DNS suffix from the robots file. Test public translation pages, licensed dictionary routes, sitemaps, media, accounts, APIs, and approved integrations separately. Pair edge controls with authentication, rate limits, signed assets, caching, data-loss monitoring, and anomaly detection.

Review checklist

Search logs for the complete registered header and preserve source IPs, ASNs, reverse-DNS results, paths, methods, response sizes, statuses, timing, and rate. Check whether the source can be independently verified and whether requested content is publicly licensed for indexing. Do not classify a request as authentic from the historical token alone.

Review the current robots file for exact product groups and record the training restriction as a content-governance requirement. Confirm that private and licensed routes are protected by application authorization and that any explicit no-training language is reflected in terms, policy, and access controls rather than robots.txt alone.

Re-check /bot, /bot.html, the current Linguee operator documentation, and the community registry when a credible accessible source becomes available. Keep this profile at documented-limit until the bot identity and current operator policy are independently confirmed. Treat the 404 as a documentation limitation, not proof of inactivity.

If you need to stop or narrow crawl, publish an exact robots group only after confirming the observed token, then verify subsequent access logs and response behavior. Do not claim that a robots rule erased previously indexed data or controlled downstream training.

References

  1. Linguee bot URL — registry-linked source; headless browser returned 404 Not Found on 2026-08-25.
  2. Linguee robots.txt — official current file with a prominent machine-translation training restriction and extensive named-agent/path rules; reviewed with the headless browser on 2026-08-25.
  3. Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
  4. RFC 9309 — Robots Exclusion Protocol standard; it does not authenticate a User-Agent.

Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.