MCP server for retrieving data from different knowledge bases with support for dense and hybrid search, embeddings, and document ingestion
This MCP server demonstrates solid definition quality with well-structured tool schemas and comprehensive parameter descriptions. All 9 tools have explicit descriptions and detailed input schemas. The naming follows verb_noun patterns consistently (retrieve_knowledge, ask_knowledge, list_*, delete_*, etc.). Parameter descriptions are detailed with constraints, enums, and format guidance. However, there are notable gaps: (1) output schemas are not documented in the tool definitions themselves, we can infer structure from descriptions but cannot see formal schema declarations for return types; (2) error handling guidance is not visible in tool descriptions, no recovery hints or categorization; (3) tool descriptions, while present, average ~150-200 chars, which is acceptable but could be more concise and action-oriented for LLM parsing; (4) security considerations (e.g., permissions, audit requirements) are not explicitly declared in tool metadata. The server excels at parameter constraints (enums, min/max, regex patterns), composition (tools are single-responsibility), and natural-language identifier acceptance (knowledge_base_name as optional fallback). Overall, this is a well-engineered knowledge base API wrapped for MCP, but lacks the LLM-optimization and error recovery patterns of top-tier tools.
Add or update a document in a knowledge base. Triggers chunking, embedding, and index update.
Search a knowledge base and use an LLM to generate an answer based on the retrieved context.
Remove a document from a knowledge base and purge its chunks from the search index.
Compare retrieval results between two index versions using the same set of queries.
Retrieve statistics about a knowledge base including chunk counts, file counts, and index state.
List all available knowledge bases with their locations and metadata.
List all registered embedding models available for retrieval and ingestion operations.
Output schemas not documented. Tool descriptions infer return structure but no formal schema is visible for retrieval results, answer responses, or stats objects. LLMs cannot plan downstream operations without knowing what fields to extract.
Error handling guidance missing. Tool descriptions do not explain failure modes, recovery steps, or error categorization (retryable vs. fatal). No guidance like 'If threshold is too high, try lowering it or use hybrid mode'.
Destructive operations lack confirmation/dry-run. delete_document and reindex_knowledge_base (with force=true) are irreversible but offer no confirmation step or dry-run mode to prevent accidental data loss.
Inferred effective spec: 2025-06-18+.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 65 | 2025-06-18+ | v2 |
| 2026-03-09 | F | 40 | - | v1 |
Trigger a full reindex of a knowledge base, scanning files and rebuilding embeddings.
Search a knowledge base using dense embeddings or hybrid (dense + BM25) retrieval and return ranked documents with optional context neighbors.
Tool annotations missing. No readOnlyHint, destructiveHint, or idempotentHint metadata visible in tool definitions. This prevents clients from marking operations with appropriate visual warnings (red delete buttons, etc.).
Permissions not declared. No scope or permission metadata (e.g., 'read:kb', 'write:kb', 'admin:kb') visible. Agents cannot be configured with least-privilege access.
No pagination limits stated for list_knowledge_bases and list_models. If a KB system has 10,000 models, unbounded responses will exhaust context. Baseline pattern requires explicit limits and pagination support.