Express-based MCP server for prior authorization request processing. Implements a multi-agent workflow (Sensing, Planning, Orchestrator, Decision agents) that processes clinical documents, extracts data, verifies member eligibility, searches coverage policies, and makes approval decisions using registered MCP tools.
This prior authorization MCP server has moderate definition quality with mixed execution. All 6 tools have descriptions and input schemas with reasonable parameter documentation. However, descriptions are often verbose and lack the concise LLM-optimized structure (target: 50-200 chars; baselines show avg 194 chars). Parameter descriptions exist but some are generic. Output schemas are not documented. Error handling and recovery guidance are absent. Tool naming uses descriptive phrases rather than verb_noun pattern (e.g., 'intelligent-document-processing' instead of 'extract_document'). No tool demonstrates idempotency hints, destructive operation warnings, or confirmation patterns despite clinical sensitivity. Composition is reasonable, tools chain logically through the PA workflow, but lack of output schema documentation means chaining must be inferred by the LLM.
Retrieves and analyzes member claims history and utilization patterns for a specified lookback period. Returns related procedures, approval rates, and billed amounts.
NLP extraction tool that identifies and validates clinical entities from documents, including patient demographics, procedure codes, diagnosis codes, and clinical findings from SOAP notes.
IDP tool extracts text, entities, and form fields from clinical PDF documents. Performs OCR, text extraction, and entity recognition on prior authorization request documents.
Verifies member eligibility, coverage status, plan type, and benefits from Member 360 data product. Returns active status, plan information, and any coverage issues.
Searches coverage policies (NCD/LCD guidelines) matching procedure and diagnosis codes. Returns matching policies, medical necessity criteria, and required documentation.
Matches clinical evidence against policy criteria. Evaluates patient information, claims history, and clinical findings to determine if policy requirements are met and recommends approval/denial/review decision.
Tool naming does not follow verb_noun pattern. Names like 'intelligent-document-processing' are descriptive but opaque to LLM intent inference. Should use verb_noun: 'extract_document', 'lookup_member_eligibility', 'search_coverage_policies', etc. Pattern baseline: 90% of A+ tools start with action verb.
Output schemas are not documented. Tool descriptions state what is returned (e.g., 'Returns member plan information, effective dates, coverage levels') but no structured schema is provided in source code. LLMs cannot see what fields to expect, forcing inference from vague descriptions. Without documented output, multi-tool chaining lacks clarity.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 6 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 53 | - | v1 |
No error handling or recovery guidance visible. Tools return success/failure but offer no actionable error messages or next steps. E.g., if 'member-eligibility-lookup' fails with 'Member not found', the tool should suggest trying different identifiers or returning available alternatives. Pattern baseline: A+ tools categorize errors as retryable/user-fixable/fatal.
Clinical-sensitive operations (like document processing and eligibility verification) lack confirmation/dry-run patterns. In healthcare PA workflows, failures can cascade, agents should be able to preview results before committing. No evidence of idempotency hints or destructive operation warnings.
Parameter descriptions exist but some are generic and lack actionable constraints. E.g., 'documentType': 'Type of document (e.g., prior-auth-request)', the 'e.g.' signals example-based documentation, but no enum is provided. LLMs may hallucinate document types. Similarly, 'requestContext' and 'clinicalEvidence' are described as objects with no nested schema detail.
Tool descriptions are verbose (many >200 chars). Baseline avg is 194 chars; some descriptions here exceed 250. E.g., 'intelligent-document-processing' description is 181 chars, acceptable but at upper boundary. LLM-optimized descriptions should be concise (50-200 chars) with WHAT/WHEN/RESULT structure.
No pagination guidance for list-returning tools. 'claims-history-retrieval' returns claims matching criteria but no limit, offset, or pagination is documented. If a member has years of claims, returning all at once bloats context and risks hallucination. Add page_size, offset, and total_count to output.
Parameter relationships not documented. E.g., 'clinical-data-extraction' has 'extractionType' enum (patient-demographics, medical-codes, clinical-findings, all) but no guidance on when 'documentData' vs 'rawText' is preferred. If both are required, say so. If one is sufficient, document that. Undocumented dependencies cause silent misuse.