Multiple MCP/ACP servers for RAG-based document querying, health information, insurance policy coverage, and doctor discovery. Includes CrewAI agents, smolagents CodeAgent, and FastMCP implementations.
Three tools across mixed frameworks (acp-sdk, FastMCP, smolagents). Critical issues: tool definitions are inferred from decorator registrations rather than explicit schema definitions; input schemas are present but poorly structured; descriptions lack clarity and specificity for LLM decision-making; no output schema documentation; parameters underspecified; no error handling guidance. The codebase shows multiple framework incompatibilities (acp-sdk agents vs FastMCP server) suggesting incomplete integration. No evidence of permissions, audit trails, or security controls.
This is a CodeAgent which supports the hospital to handle health based questions for patients. Current or prospective patients can use it to find answers about their health and hospital treatments.
This tool returns doctors that may be near you.
This is an agent for questions around policy coverage, it uses a RAG pattern to find answers based on policy documentation. Use it to help answer questions on coverage and waiting periods.
Input schemas lack proper structure and type documentation. policy_agent and health_agent accept 'input: list[Message]', the Message type is not defined in the schema, forcing LLMs to guess at the structure. list_doctors accepts 'state: string' with no enum, pattern, length, or format constraints.
No output schema documented for any tool. LLMs cannot predict what fields will be returned, forcing them to reason about downstream tool compatibility without guidance. policy_agent returns 'Message(parts=[MessagePart(content=str(task_output))])', structure is implicit, not documented.
Descriptions are vague and lack actionable context. 'This tool returns doctors that may be near you' (49 chars) does not explain when to call it vs. a general search, what 'near you' means, or what data is returned. 'This is an agent for questions around policy coverage' (54 chars) lacks specificity, what policy? What coverage types? Any prerequisites?
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 35 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 30 | - | v1 |
Tool definitions are inferred from decorators (@server.agent(), @mcp.tool()) rather than explicit registration with full schemas. Code shows files are separate (crew_agent_server.py, mcpserver.py, smolagents_server.py) with no unified tool manifest. Per HARD SCORING RULE: 'If you cannot see the actual tool definition in the source (only inferred): cap that tool's overall at 50.' All three tools hit this cap.
No parameter descriptions for complex inputs. health_agent accepts 'context: Context' with no description of what Context contains or why it's needed. list_doctors 'state' parameter has no format constraint (expecting 2-letter ISO code, but no validation documented).
No error handling or recovery guidance. policy_agent calls crew.kickoff_async() with no try-catch or error classification. list_doctors fetches from GitHub URL with no timeout, retry logic, or guidance on what to do if the request fails. No tool tells the LLM 'if this fails, try X' or whether errors are retryable.
No validation of required parameters. list_doctors accepts state as string with no validation that it's a valid 2-letter code. policy_agent and health_agent assume input[0].parts[0].content exists but do not validate array bounds or part presence, risking IndexError.
Framework fragmentation. Three separate server files (crew_agent_server.py, mcpserver.py, smolagents_server.py) use different frameworks (acp-sdk, FastMCP, smolagents) with inconsistent tool registration patterns. No unified tool discovery or coordination. Unclear how agents are deployed together or if they can be called from the same client.
No security controls. policy_agent loads PDFs from a hardcoded relative path '../data/gold-hospital-and-premium-extras.pdf' without validation. health_agent calls external tools (DuckDuckGo, webpage visit) without permission gates or scope declarations. No audit trail, no permission checks, no secrets management.
Naming lacks verb-noun clarity. 'policy_agent' and 'health_agent' are nouns, not action verbs. Better names would be 'answer_policy_question', 'search_health_info', 'query_hospital_policy'. 'list_doctors' is acceptable but generic, 'search_nearby_doctors' or 'find_doctors_by_state' is clearer.
list_doctors returns raw JSON string conversion of doctor objects. No output schema documented. LLM must parse string output to extract fields. Response includes all fields from raw JSON (address, specialty, etc.) but no clarity on which fields are guaranteed or what the structure is.
No pagination support. list_doctors loads entire doctor JSON from GitHub and filters in-memory. If the list grows large, all doctors are fetched and returned, no limit enforced, no offset/page parameters, no total count returned.