A healthcare AI backend system with RAG (Retrieval-Augmented Generation) capabilities, medical knowledge management, and LangGraph-based workflow orchestration for clinical consultations
This HTTP-based medical AI backend exposes 8 tools with significant quality gaps. Tool naming follows verb patterns adequately (search_, lookup_, create_, upload_, read_), but descriptions are inconsistent in quality and depth. Most critically, input schemas are partially visible but output schemas are entirely undocumented, and parameter descriptions lack the specificity required for production agent use. The server implements medical RAG and patient lookup tools that would benefit from explicit error handling guidance for sensitive healthcare operations.
Endpoint for baseline Naive RAG mode. Provides a simpler retrieval-augmented generation without the full workflow orchestration.
Endpoint for conversing with the main workflow. Receives user messages and routes them through the LangGraph workflow for processing with medical expertise.
Creates and vectorizes a new document manually, storing it in the knowledge base for RAG retrieval.
Health check endpoint that returns server status.
Tool for the patient worker node to look up patient clinical history and records from the database.
Retrieves all documents from the knowledge base.
Endpoint for semantic search across the knowledge base. Searches for N most similar document chunks to a query.
Search tool for RAG that queries the knowledge base to retrieve relevant documents. Bound to the medical agent node for context retrieval during agent execution.
Output schemas completely undocumented. No visible documentation of what fields each tool returns, their types, or structure. LLMs cannot plan downstream tool calls or extract required data.
No error handling documentation. Tools lack guidance on failure modes, recovery paths, or what the LLM should do if a call fails. Critical for sensitive healthcare operations (patient lookup, knowledge creation).
Parameter descriptions lack actionable constraints. 'query' parameters do not specify format, length limits, or valid input patterns. No enum constraints for 'model' parameter in chat tool despite being a controlled set.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 11 | <=2025-11-25 | v2 |
| 2026-03-09 | D | 56 | - | v1 |
Uploads a PDF file, divides it into chunks, and stores it in the vector database for RAG.
Naming ambiguity: 'search_knowledge_base', 'search_knowledge', and 'read_knowledge' are confusingly similar. The distinction between them is unclear, which should the LLM use? read_knowledge returns all docs; search_knowledge searches by similarity; search_knowledge_base is a duplicate concept.
read_knowledge has no input schema (empty {}), but no documentation explaining what pagination it supports or how many documents it returns. Returning 'all documents' without pagination will blow context windows.
upload_pdf uses generic 'file' type parameter. No description of supported formats, file size limits, chunk strategy, or expected vectorization process. Medical PDFs may contain sensitive data, no security or privacy documentation.
'chat' and 'chat_naive' both appear to accept queries but have different underlying implementations (LangGraph agent vs naive RAG). No clear guidance on when to use which. No retry/idempotency guidance for conversation state management.
Patient lookup tool (lookup_patient_history) lacks any documentation of privacy/security implications, data retention, or permission checks. Healthcare data access must be audited, no logging or security annotations visible.
search_knowledge parameter 'limit' has type integer with no min/max bounds documented. LLM could pass 0, negative, or million-sized limits. Baseline rubric shows numeric params need explicit ranges (1-100, 1-365, etc.).