MCP server for clinical risk score calculation using LLM-based information extraction and validated scoring algorithms (CHA₂DS₂-VASc, EuroSCORE II, HAS-BLED)
The MCP Risk Server exposes a single tool 'llm_pipeline' with a clear description and basic parameter schema. However, the tool has a critical security concern (file system access via parameter), lacks error handling guidance, has no documented output schema, and the parameter descriptions are minimal. The server demonstrates foundational structure but falls significantly short of production quality standards. Tool definitions exist and are registered, but implementation details show poor defensive practices around file I/O.
Executes a two-stage risk score calculation pipeline: Stage 1 extracts clinical information from medical texts using an LLM, Stage 2 calculates validated risk scores (CHA₂DS₂-VASc, EuroSCORE II, or HAS-BLED) from the extracted data
No output schema documented. Tool description explains what it does internally (two-stage pipeline, risk score calculation) but does not specify what the agent receives back. Without output schema documentation, LLMs cannot plan downstream actions or extract key fields (e.g., risk score value, confidence, extracted clinical data fields).
Parameter descriptions are minimal (under 20 characters for 'data_folder' and 'config_file'). 'Path to directory containing input text files with medical data' lacks actionable format guidance, does it accept relative or absolute paths? What file formats? Are nested directories recursively processed? These ambiguities force LLMs to guess and risk invalid inputs.
File system paths exposed as tool parameters represent a critical security vulnerability (pattern:secret-injection violation). The tool accepts 'data_folder' and 'config_file' parameters, but there is no evidence of path sanitization or access control in the visible code. An LLM could be tricked into reading arbitrary files (e.g., /etc/passwd, .env files with credentials) or writing to sensitive directories. This violates the principle of least privilege and audit-trail requirements.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 52 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 30 | - | v1 |
No error recovery guidance. Tool description does not indicate what happens on failure: Does the tool return partial results? Does it fail the entire pipeline if Stage 1 or Stage 2 encounters an error? Are there retry semantics? An LLM has no guidance on whether to retry, ask the user, or abort downstream steps.
Tool naming is non-standard. 'llm_pipeline' uses snake_case (acceptable) but does not follow the verb_noun convention (e.g., 'execute_risk_pipeline', 'calculate_risk_score', 'process_medical_data'). The name does not clearly signal what action the agent triggers or what state change occurs (two-stage LLM pipeline and risk calculation are side effects, not obvious from the name alone).
No input validation or constraint documentation. Parameters accept arbitrary file paths and config file references, but there is no schema-level validation (e.g., pattern, minLength, maxLength) or description-level guidance on allowed directory structures, config file format requirements (YAML validation rules), or maximum file sizes to prevent DOS attacks.
No indication of idempotency or side effects. The tool description says it 'calculates' risk scores, but does it write files? Does it modify configuration? Can an agent safely retry if interrupted? This ambiguity violates the idempotent-operation pattern and risks duplicate work (e.g., duplicate output files, repeated LLM calls).