MCP server for banking compliance: KYC/AML, sanctions screening, and regulatory reporting
Banking Compliance Helper has 5 tools with complete input schemas and descriptions. All tools have proper Pydantic models with typed fields and descriptions. However, output schemas are not formally documented, descriptions lack nuance about when to use tools vs. alternatives, and error handling is minimal. Tool naming is strong (verb_noun pattern: check_, screen_, assess_, generate_, query_). Parameters are well-typed with enums where appropriate (e.g., sender_type, subject_type). The server avoids secrets in parameters. However, the descriptions are functional but brief (100-200 chars), missing guidance on dependencies between tools and recovery paths. No parameter-level validation constraints stated (e.g., amount ranges, date formats). Output is structured (dict/JSON) but lacks documented field definitions for downstream tool chaining.
Assess Anti-Money Laundering (AML) risk for a transaction. Analyzes red flags including amount, jurisdictions, transaction type, and patterns indicative of money laundering or terrorist financing.
Screen individual or entity against OFAC, EU, and UN sanctions lists. Returns risk assessment and match probability based on provided information. Real implementation would query actual sanctions databases.
Generate a draft Suspicious Activity Report (SAR) for filing with FinCEN. Creates structured SAR narrative based on detected suspicious activity. Output should be reviewed by compliance officer before filing.
Look up key requirements from banking regulations (Basel III, GDPR, PSD2, etc). Provides quick reference to common regulatory requirements for banking compliance.
Check if individual is a Politically Exposed Person (PEP). Screens against global PEP databases for current/former government officials, senior executives of state-owned enterprises, and their close associates.
Output schemas not documented. Tools return dicts with fields like 'risk_level', 'risk_score', 'recommendation', 'risk_factors', etc., but there is no formal schema definition for downstream tool selection or field chaining. LLMs cannot predict what fields are available, forcing them to hallucinate or make unsafe assumptions.
Parameter descriptions lack format and constraint guidance. For example, 'dob' is described only as 'Date of birth (YYYY-MM-DD)' but does not state min/max valid ranges, whether historical dates only are allowed, or what error is returned for invalid formats. 'amount' in assess_transaction_risk has no min/max declared; LLMs may pass negative or zero values.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 69 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 43 | - | v1 |
Tool descriptions are functional but lack context on WHEN to use each tool and what the agent should do with the output. For example, check_sanctions and screen_pep both return 'recommendation' fields, is the agent expected to act on these? What does 'REVIEW' or 'BLOCK' mean in terms of downstream actions? The description should state: 'Returns a recommendation (BLOCK/REVIEW/PROCEED) that the agent should use as input to generate_sar_report or escalate to a compliance officer.'
Error handling is missing. Functions return success dicts even when validation fails. For example, if 'dob' is malformed, there is no validation error, the function still returns a result. LLMs need explicit error messages like 'Invalid date of birth format: expected YYYY-MM-DD, got "invalid". Please retry with a valid date.' to self-correct.
generate_sar_report is marked as a WRITE tool (generates a report for filing), but there is no dry-run or confirmation step. Agents may call this without human review, creating a false SAR filing. Compliance tools must gate destructive operations behind confirmation. Add a 'draft_only' parameter (default true) so agents generate drafts first, and require explicit 'draft_only=false' + confirmation before filing.
No tool composition guidance. The server has 5 tools but does not document how they work together. For example, after check_sanctions returns 'HIGH' risk, should the agent call assess_transaction_risk? After screen_pep returns is_pep=true, what next? The tool descriptions should include: 'Typically called before assess_transaction_risk or generate_sar_report to gate compliance decisions.'
query_regulation is vague. It accepts a 'regulation' parameter like 'basel iii' or 'gdpr' but does not enumerate valid values or describe what the returned 'requirements' field contains. Is it a list of URLs, a prose summary, a structured object with key/value pairs? The description should specify: 'Returns a dict with fields: regulation_name, short_summary (200 chars), key_requirements (list of strings), applicable_jurisdictions (list), enforcement_body.'