A multi-agent compliance checking system using MCP, LangChain, and RAG for financial risk assessment and regulatory compliance verification.
This MCP server exhibits significant definition quality issues across multiple dimensions. While all 6 tools have basic descriptions and input schemas, most descriptions are generic or underdescriptive (<100 chars), and parameter annotations lack critical details. The server mixes infrastructure concerns (RAG indexing, LLM generation) with business logic (risk reporting, compliance querying) in a way that violates single-responsibility principles. Tool names are mostly verb-starting but lack clarity about their prerequisites and interconnections. Error handling is minimal, most tools return generic error messages without recovery guidance. Output schemas are largely undocumented, forcing LLMs to infer result structure. The server was tested for HTTP transport capability: it appears to be Flask-based HTTP, which is positive, but MCP registration/discovery endpoints are not visible in the provided code, making tool definitions potentially inferred rather than explicitly registered.
A2A endpoint to generate risk and compliance report.
A2A endpoint to provide the selected client profile.
Generate text using ChatOpenAI.
Index regulation files into the vector database.
Query the RAG service.
A2A endpoint to receive risk report and query compliance actions.
Output schemas are undocumented. The source code shows that tools return results (e.g., rag_query returns {context, result, citations}; generate_risk_report returns structured data) but no explicit schema definitions are provided in tool metadata. LLMs cannot infer downstream field names and types, forcing them to guess and risking failed downstream tool chaining.
Tool descriptions are generic and lack 'WHEN to use' guidance. For example, 'Query the RAG service' (rag_query) does not explain when to call rag_query vs openai_generate, or how they differ functionally. Descriptions under 100 chars fail to disambiguate similar tools.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 47 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 42 | - | v1 |
Parameter descriptions are minimal or missing context. The 'client' parameter in rag_query is described as 'Client information object' but lacks detail: what fields must it contain? What is its structure? Is it required? Similar issues affect 'context' (openai_generate) and 'client_info' (generate_risk_report).
No input validation constraints or error guidance. Parameters like 'chunk_size' (rag_index) and 'risk_level' (receive_risk_report) lack min/max bounds or enum constraints. If an LLM passes chunk_size=999999, there is no documented maximum. 'risk_level' should be an enum (low/medium/high), not a free-form string.
Error handling is absent or generic. The code shows error returns like {'context': 'Error retrieving context', 'result': 'Error querying RAG: ...', 'citations': []} but does not guide the LLM on recovery. No distinction between retryable errors (temporary service outage) and user-fixable errors (file not found, invalid query). LLMs need explicit error classification to know whether to retry, ask the user, or abort.
Tool composition and single-responsibility violations. The server conflates infrastructure (rag_index, rag_query, openai_generate) with business logic (get_client, generate_risk_report, receive_risk_report). rag_index performs file loading, chunking, and embedding, three concerns. This makes it hard for agents to compose them granularly and increases failure surface.
Tool naming does not reflect prerequisites or dependencies clearly. 'receive_risk_report' suggests it receives a report from an external source, but it actually processes results from generate_risk_report and queries compliance actions. Naming should be 'query_compliance_actions_from_report' or similar to clarify the flow.
No documented pagination or result-limiting behavior. If rag_query returns 1000 documents, the context window will be exhausted. The code lacks result limits, and no limit parameter is documented. Production tools must cap results (typically 20 - 50) and return cursors for pagination.