MCP Server providing RAG capabilities for Continue IDE. Implements a Retrieval-Augmented Generation system with document ingestion, vector storage, and LLM integration.
This FastAPI-based MCP server exposes 10 RAG and LLM configuration tools via HTTP. Tool definitions are visible in mcp_server.py and api/endpoints.py with explicit Pydantic models (QueryRequest, IngestRequest, etc.), but quality is inconsistent. Strengths: most tools have descriptions and basic input schemas with types. Weaknesses: (1) duplicate tool definitions (get_collection_info appears twice, tools 3 and 7), (2) weak descriptions lacking actionable context ('Update LLM configuration' is vague, what happens if you pass an invalid model name?), (3) several tools lack descriptions entirely in the observable schema, (4) no output schemas documented beyond Pydantic response models, (5) critical security issue: api_key exposed as a parameter in update_llm_config (tool 9), (6) no error recovery guidance in descriptions, (7) tools lack hints about dependencies (e.g., query requires RAG to be initialized), (8) no per-tool idempotency or destructiveness annotations visible in the MCP registration layer. Naming is mostly verb-first and clear (query_rag, ingest_document, clear_collection), but several naming issues: ingest_document vs ingest_file vs ingest_path are functionally similar and risk LLM confusion; query and query_rag serve overlapping purposes.
Clear all documents from the collection
Get available LLM models
Get information about the vector collection
Get information about the vector collection
Ingest a document into the RAG system
Ingest a document from file upload
Ingest a document from file path
CRITICAL: api_key exposed as tool parameter in update_llm_config. Credentials must never appear in parameters, they will be logged and leaked into traces. Use server-side secret injection via environment variables.
Duplicate tool definition: get_collection_info appears twice (tools 3 and 7). This will confuse MCP clients and agents about which to call. Remove one or merge into a single canonical definition.
Three similar ingest tools (ingest_document, ingest_file, ingest_path) with overlapping semantics. LLMs cannot easily distinguish when to use each. Consolidate into a single parameterized ingest tool or clearly document the use case for each variant.
Two query tools (query_rag and query) with overlapping functionality. The descriptions do not clarify the difference. Agents will waste reasoning cycles deciding between them. Use one canonical name or document when each should be called.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 50 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 37 | - | v1 |
Query the RAG system
Query the RAG system with a question
Update LLM configuration
Tool descriptions lack actionable context and error recovery guidance. Example: 'Update LLM configuration' does not explain what happens on invalid inputs, which fields are optional, or how to recover from a failed update. Descriptions should answer: WHAT does it do? WHEN should I call it? WHAT do I get back? WHAT if it fails?
No output schemas documented in tool definitions. Response models exist in code (QueryResponse, IngestResponse, HealthResponse) but are not mapped back to the MCP tool registration layer. LLMs do not know what fields to expect in responses, forcing them to guess and parse unstructured output.
No tool annotations visible (readOnlyHint, destructiveHint, idempotentHint). The rubric notes tools as risk=DESTRUCTIVE (clear_collection) or risk=WRITE, but these are not encoded in the MCP protocol layer. Agents cannot reason about safety without explicit annotations.
Parameter descriptions are minimal or missing. Examples: top_k is described as 'Number of relevant chunks' but does not specify why the default is 5 or when to increase it; api_base in update_llm_config lacks guidance on format (URL scheme required?); temperature has range [0, 2] but no explanation of the semantic difference.
No documented error handling or recovery guidance. If ingest_document fails on a corrupted PDF, the error message does not guide the agent to retry, try a different format, or call a diagnostic tool. Error responses must be actionable, not just status codes.
clear_collection (DESTRUCTIVE tool) lacks a confirmation step. Agents should be required to explicitly confirm before irreversibly clearing all documents. A dry-run parameter or a separate confirm_collection_clear tool would prevent accidental data loss.