Multi-agent security system for evaluating prompts for threats, masking PII data, and analyzing SQL queries safely. Implements Agent-to-Agent (A2A) protocol and MCP (Model Context Protocol) with FastMCP for tool exposure.
This MCP server has critical definition quality gaps across nearly all dimensions. Tools lack proper input schemas with type information, descriptions are minimal or absent, and there is no documented output structure. The codebase shows tool definitions are inferred from agent code rather than explicitly registered via MCP mechanisms. Tool descriptions are generic and do not follow LLM-optimized patterns. Parameters lack type constraints, enums, and validation guidance. No error recovery guidance is present. The server appears to mix MCP transport (STDIO) with a separate A2A Agent framework, creating confusion about the actual tool surface. Evidence: (1) execute_sql_query and query_data both claim 'Execute SQL queries safely' but are listed as different risk levels, naming and responsibility are unclear. (2) No input schemas visible for any tool, only tool names and descriptions inferred from the description field. (3) descriptions are under 60 characters for most tools, failing to explain WHEN to use each tool or WHAT the agent should expect in return. (4) mask_text references 'Google Cloud DLP' but no documentation of required setup, credentials, or when to call it. (5) evaluator tool has no parameter description for 'text', LLMs cannot infer what security threats it detects or how it differs from other validation tools.
Evaluates prompts for security threats.
Execute SQL queries safely on the salaries database.
Get schema and sample data for specified tables (comma-separated).
List all tables in the database.
Masks sensitive data like PII in text using Google Cloud DLP.
Execute SQL queries safely on the salaries database.
No input schemas visible in source code. All 6 tools lack JSON Schema definitions with type information, constraints, and field descriptions. Only tool names and brief descriptions provided.
Duplicate tool responsibility: execute_sql_query and query_data both claim to 'Execute SQL queries safely' but are listed with different risk levels (WRITE vs READ_ONLY). Naming does not disambiguate; LLMs will conflate them.
Tool descriptions are below 60 characters, providing insufficient context for LLM tool selection. No WHEN to use, WHY to choose this tool over similar ones, or WHAT the return value contains. Descriptions do not answer: what does it do, when should I call it, what does it return?
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-21 | F | 33 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 41 | - | v1 |
Parameter descriptions are absent or generic. The 'sql' parameter in execute_sql_query and query_data has only 'SQL query to execute', no guidance on supported SQL dialects, max query length, timeout behavior, or what happens on syntax error.
No output schemas documented. Tools do not document what fields they return, data types, or structure. LLMs cannot plan downstream calls or extract required data (e.g., does list_database_tables return table names as strings, or objects with schema info?).
No error handling guidance. No documentation of what errors can occur, whether they are retryable, or what recovery actions the LLM should take. Tool descriptions do not include error classification or recovery hints.
SQL injection risk not addressed in tool definitions. execute_sql_query and query_data accept arbitrary SQL but provide no documentation of sanitization, parameterized query support, or what happens if an LLM passes malicious input. No validation rules documented.
mask_text references 'Google Cloud DLP' but does not document setup requirements, authentication, cost, or when to invoke it. No description of what data types it masks (PII, PHI, financial) or return format.
evaluator tool has minimal description ('Evaluates prompts for security threats') and 'text' parameter lacks any description. No guidance on what constitutes a security threat (SQL injection, XSS, prompt injection, etc.) or how the evaluation result is returned.
Tool definitions appear inferred from agent code rather than explicitly registered via MCP. No fastmcp.Tool decorators visible with schemas. Tool names and descriptions extracted from docstrings or comments, not formal registrations. This prevents confident assessment of actual parameter types.