A Documentation RAG MCP server for semantic search and retrieval of documentation using embeddings and reranking
This server exposes 3 tools with inconsistent quality. Tool schemas are present but parameter descriptions are sparse or missing entirely. The server demonstrates basic RAG functionality but lacks proper error handling guidance, output schema documentation, and comprehensive parameter validation. Naming conventions are reasonable (verb-first: ask-, get-), but descriptions are generic and do not guide LLM selection effectively. This is typical of an early-stage community RAG server; it works for basic use cases but falls short of production-grade agent integration.
Ask a question and get a synthesized answer using the RAG agent and contextually relevant information from the embedding database. Optionally filter or boost by tags.
Retrieves relevant documentation to provide Cline with context for project tasks. Optionally filter or boost by tags.
Generate an embedding vector for the given input text using the Ollama API.
Parameter descriptions missing or incomplete. 'query' parameter in ask-question and get-context lack descriptions explaining what kind of questions or queries yield best results. 'tags' parameter description is present but terse ('Optional list of tags to filter or boost relevant chunks'), does not explain how tag filtering works, what tags exist, or when to use it. 'model' parameter in get-embedding has no description.
No output schema documented for any tool. ask-question and get-context return search results and synthesis, but the LLM does not know the structure (are they lists of objects? each with what fields?). get-embedding returns a vector, but vector dimensionality, format, and type are undocumented. LLMs cannot plan downstream calls or extract specific fields without this schema.
No error handling or recovery guidance. Tool descriptions do not mention failure modes, when to retry, what to do if the embedding service is unavailable, or if a query returns zero results. LLMs have no fallback strategy.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 44 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 13 | - | v1 |
'tags' parameter in ask-question and get-context is under-specified. No enum values, no description of valid tags, no guidance on tag semantics (are they filter-only or do they boost results?). LLM must guess valid values, risking silent failures.
'top_n' parameter has a description but no bounds (minimum/maximum). Unbounded integer allows LLM to pass absurd values (e.g. 999999) which may break the API or cause timeouts.
'get-embedding' description is generic ('Generate an embedding vector...'), does not explain when to call it vs. ask-question or get-context, what the embedding is used for, or what the vector dimensions are. LLM may conflate it with the other retrieval tools.