Poggio-Citations: Robots.txt & Crawl Policy Reference
Technical reference for the Poggio-Citations User-Agent label. Learn how to investigate citation-oriented traffic when the linked crawler endpoint is unavailable.
AI Summary:
Poggio-Citations 1.0is a registry label associated with Poggio's enterprise AI revenue-intelligence platform, but the linked crawler endpoint currently redirects to a first-party 404 page and the current documentation does not publish a crawler contract for this token. Treat the header as an observation, not authenticated identity; verify traffic from logs before applying a robots or WAF policy.
Role and policy boundary
Poggio's current official documentation describes an enterprise revenue-intelligence platform with a Deep Research Agent that monitors web, news, financial filings, and social signals. It also documents a unified context engine, MCP server, REST API, and on-demand revenue workflows. This establishes that Poggio has web-research capabilities, but it does not prove that a request carrying Poggio-Citations 1.0 is currently operated by Poggio or that it is a training-data crawler.
The registry links https://docs.poggio.io/api/robots as the bot documentation. In this review, that address redirected to https://poggio.io/api/robots and returned a Poggio 404 page. The current docs at https://poggio.io/docs describe the product but do not publish this User-Agent, source ranges, crawl rate, verification procedure, or robots behavior.
Treat Poggio-Citations as a legacy or partially documented label. Blocking it may affect one possible research or citation path, but it does not control ordinary search engines, other AI systems, or Poggio clients using another identity. Allowing it does not grant access to private customer, paywalled, or authenticated content.
If your logs confirm the exact token and you want to communicate a restriction, publish:
User-agent: Poggio-Citations
Disallow: /
For a selective public-reference policy:
User-agent: Poggio-Citations
Allow: /docs/
Allow: /public-reference/
Disallow: /private/
Disallow: /account/
Disallow: /internal/
Disallow: /api/
Robots.txt is advisory. Use authentication and authorization for confidential content.
Layered verification
Start with raw access logs and preserve the complete User-Agent, source IP, ASN, reverse DNS, method, path, status, response size, redirects, timestamp, and request rate. The registry token is self-declared, and no current Poggio source range or reverse-DNS verification procedure was confirmed.
Analyze behavior without over-assigning purpose. A research workflow may request individual news or company pages, while an automated monitor may visit many domains, feeds, sitemaps, and financial pages. These patterns can explain operational impact but cannot prove that a request is a citation fetch, indexing job, or training collection.
Evaluate /robots.txt independently. Confirm the canonical host, response status, content type, exact group, and path match. Because the linked policy endpoint is currently a 404, do not infer Poggio's own compliance behavior from its absence. For page-level discoverability, inspect:
<meta name="robots" content="noindex, nofollow">
X-Robots-Tag: noindex, nofollow
These directives do not provide confidentiality and may not stop an undocumented or user-triggered client. Protect private pages and APIs at the application layer.
WAF and Nginx remediation examples
If logs confirm unwanted requests carrying the token, a narrow WAF match can block the declared identity:
{
"description": "Block observed Poggio-Citations token",
"expression": "lower(http.user_agent) contains \"poggio-citations\"",
"action": "block"
}
For Nginx, scope enforcement to sensitive routes while investigating public research access:
map $http_user_agent $block_poggio_citations {
default 0;
~*Poggio-Citations 1;
}
server {
location ~ ^/(private|internal|account|paywall|api)/ {
if ($block_poggio_citations) { return 403; }
try_files $uri $uri/ =404;
}
}
A header rule is easy to spoof or evade and may block a legitimate approved workflow. Test ordinary browsers, search crawlers, social previews, internal monitors, and any approved Poggio integration. Combine the rule with rate limits, route authorization, signed URLs, and anomaly detection rather than treating the token as proof of origin.
Review checklist
Search logs for the exact Poggio-Citations value and preserve source network, request paths, response sizes, statuses, and timing. Re-check the registry-linked endpoint and current Poggio docs; during this review the former returned a first-party 404 and the latter documented the product without a crawler policy.
Decide whether your objective is to preserve citation-oriented visibility, prevent possible extraction, protect private sales intelligence, or reduce crawl load. Publish a targeted robots group for the exact observed token, enforce sensitive routes with WAF and application authorization, and test public docs, news pages, feeds, sitemaps, APIs, and authenticated content separately. Revisit this profile if Poggio publishes a working crawler page, canonical User-Agent, source verification method, or explicit robots policy.
References
- Poggio crawler URL — registry-linked endpoint; it redirected to a first-party 404 during the 2026-08-24 review.
- Poggio Docs Overview — official product documentation describing web/news monitoring and research capabilities, but not this crawler token.
- Google Robots.txt Introduction — general explanation of crawler directives and their limitations.
Need to optimize your entire site for AI search visibility? Run a comprehensive audit with Geolify.ai.