A multi-service RAG (Retrieval-Augmented Generation) system with MCP redaction gate, integrating OpenAI LLM, Qdrant vector store, and PII/secrets redaction.
This MCP server exposes 4 tools with mixed quality. Two tools (safe_send_webhook, safe_store_memory) are defined in ai_service with Pydantic schemas and descriptions, but two tools (redact_text, classify_sensitivity) are registered via FastMCP in redaction_gate with minimal documentation. Tool naming is reasonably clear (verb_noun pattern), but descriptions lack implementation details and error guidance. Schemas are present for safe_* tools but parameter descriptions are inconsistent. The redaction_gate tools lack visible input schemas in the provided code. No tool includes error handling guidance or recovery paths. Output schemas are not documented anywhere. This is a typical community MCP server, competent but incomplete.
Classify sensitivity level (low/medium/high) based on detected PII/secrets.
Redact common PII/secrets from text before storing/sending.
Send text to an approved external webhook destination in a safe way. This tool ALWAYS redacts PII and secrets from the text before sending. Use this when the user explicitly asks to notify, post, or send information to an external system (e.g. Slack, Teams, internal webhooks). The destination must be on the approved allowlist. Do NOT use this tool for storage or internal logging.
Store long-lived information or preferences for future conversations. This tool ALWAYS redacts PII and secrets before storing data. Use this only for durable facts the user wants remembered (e.g. preferences, project context, decisions), not for transient chat logs. Choose an appropriate scope (user, tenant, or session). Do NOT store sensitive identifiers, credentials, or private data.
redact_text and classify_sensitivity lack visible input schemas in source code. The redaction_gate/app/tools/redact_tools.py shows @mcp_server.tool() decorator with no args_schema parameter, and payload is received as bare dict without schema validation. This violates the requirement that all tools expose formal JSON Schema.
classify_sensitivity has a 20-character description ('Classify sensitivity level (low/medium/high) based on detected PII/secrets.') which is below the 34-char p10 baseline. It does not explain WHEN to call it vs redact_text, what the return format is, or how the classification affects downstream behavior. This violates the minimum expectation for tool discoverability.
No output schemas are documented for any tool. LLMs cannot plan downstream tool calls or extract required data (e.g., what fields does classify_sensitivity return? is it 'sensitivity_level' or 'level'?). This blocks tool composition and forces LLMs to guess field names.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 46 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 49 | - | v1 |
safe_send_webhook and safe_store_memory return responses like {'ok': bool, 'status_code': int} and {'ok': bool, 'stored_key': str, 'scope': str} but these are returned as plain dicts in async functions with no type hints. The actual output structure is not declared in docstrings or schemas, making it impossible for downstream tools to rely on specific fields.
Error handling is minimal. safe_send_webhook returns {'ok': False, 'error': '...'} for allowlist violations, but does not suggest recovery steps (e.g., 'Contact admin to add this host to ALLOWED_WEBHOOK_HOSTS'). No error categorization (retryable vs. user-fixable vs. fatal). This prevents agents from reasoning about recovery paths.
The 'actor' parameter (Dict[str, Any]) is accepted in all tools but never described in parameter annotations. LLMs cannot infer what fields are required: is it {'user_id': str, 'tenant_id': str}? What if one is missing? This forces agents to guess or log errors.
The 'scope' parameter in safe_store_memory says it accepts 'user, tenant, or session' but is typed as a bare string with no enum constraint or validation. LLMs can pass invalid values like 'group' or 'org', and the tool will silently accept them.