MCP Proxy Gateway for AgentGuard — runtime safety layer for AI agents with cryptographic auditing, real-time control, and forensic tracing
Scoring was not performed
All tool descriptions are under 50 characters, well below the 50-200 char LLM-optimized baseline. Descriptions lack WHEN to use the tool and WHAT it returns, forcing LLMs to guess at selection logic and output structure.
No output schemas documented. LLMs cannot reason about what fields to expect or plan downstream tool calls. For example, query_traces likely returns traces with agent_id, timestamp, action fields, but this is not stated.
query_traces and list_violations accept optional 'limit' parameter but do not document maximum limits (e.g. 'Max 100') in the description. Input description says 'Max results (default 20, max 100)' for query_traces but this constraint is not enforced schema-side (no maximum in JSON Schema).
No error handling guidance. If a tool call fails (e.g. agent_id not found), the LLM receives no recovery hint. Responses should guide: 'Agent not found. Try list_agents() first or search by name.'
list_policies accepts no parameters and has no description of the expected output structure (how many policies, what fields per policy, pagination). This invites LLM misuse.
No tool declares idempotency guarantees. All four tools are read-only (safe to retry), but this is not formally documented via tool annotations or description callouts. An LLM retrying on ambiguous network errors should be confident it's safe.
query_traces and list_violations support pagination via 'limit' but no 'offset' or 'cursor' parameter. Cannot retrieve results beyond the first page. list_* tools typically need pagination support.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 0 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 44 | 2024-11-05+ | v1 |