Semantic search engine for markdown documents. MCP server with multi-provider embeddings and Milvus vector storage.
This server has 6 tools with clear naming (verb_noun convention) and meaningful descriptions (100-200 chars typical). However, the implementation suffers from critical gaps: input schemas are NOT fully visible in the provided source code excerpt (truncated at embedding vector code), parameter descriptions are present in schema but lack depth/constraints, and output schemas are not documented. The tools address a coherent RAG domain but lack production-grade error handling, validation guidance, and per-parameter constraint documentation. STDIO transport further limits the score. The server sits in the 'Fair/Poor' range due to incomplete schema visibility and lack of structured output documentation.
Clear all indexed documents from the vector database.
Get list of background indexing jobs and their status.
Get the current status of the indexed documents collection.
Index markdown documents from a directory into the vector database for semantic search.
List all markdown files that have been indexed.
Perform semantic search over indexed markdown documents using vector similarity.
Output schemas not documented. No tool response format is specified, LLMs cannot plan downstream calls or extract required fields. search_documents likely returns SearchResult objects but callers cannot know the field structure without inspecting utils.py.
Missing parameter constraints and validation guidance. 'recursive' and 'force' boolean parameters lack enum/const documentation. 'top_k' integer has no min/max bounds stated in descriptions (should enforce 1 - 100 or similar). 'directory' path parameter lacks sanitization warnings for path traversal.
clear_index requires a 'confirm' boolean but provides no guidance on what happens if confirm=false. Should describe the consequence: operation aborted vs partial execution vs error. Current description assumes users know the semantics.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 56 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
No error handling or recovery guidance in tool descriptions. If index_documents fails due to invalid markdown or Milvus connection timeout, the LLM receives no actionable error message or retry strategy. Descriptions do not hint at prerequisites (e.g. 'Milvus must be running').
Tool descriptions lack dependency hints. search_documents does not state 'You must call index_documents first to populate the vector database.' index_documents does not document when to use 'force=true' or what happens to existing indexed data with 'recursive=true'.
No pagination support. list_indexed_files and get_background_jobs lack offset/limit/next_cursor parameters. If an index contains thousands of files, returning all of them in a single response bloats the context window and breaks LLM reasoning.
STDIO transport only. The server is not remotely accessible, cannot be used by hosted MCP clients or deployed as a shared service. Hard-caps protocol readiness and production deployment options.