MCP server for hybrid document search combining vector search and BM25 full-text search with support for multiple embedding providers and languages
The server defines three tools with explicit schemas and descriptions visible in server/src/mcp/tools.rs. Naming is verb-based (search, get, get_project_info) and clear. Descriptions exist for all tools and are substantive (60-150 chars each). However, several parameters lack descriptions (e.g., top_k is described as 'Number of results to return (default: 10)' but no range bounds; filters object has no description; source_type and path_prefix lack context for when to use them). Output schemas are not documented, the code shows ToolResult with content/isError fields, but tool descriptions don't explain what search() returns (result structure, fields, ranking info). Error handling exists (ToolResult.error()) but descriptions lack recovery guidance. Parameters lack validation constraints (e.g., top_k type is 'number' but no min/max; source_type is free-form string instead of enum with allowed values like ['md','txt','pdf','xlsx']). The schema quality is decent but incomplete for production use.
Get the full content of a specific document chunk by its chunk_id.
Get information about the current project: collection name, document count, tantivy index directory, and embedding settings.
Search documents using hybrid search (vector + BM25). Returns ranked results from indexed documents.
search tool: top_k parameter lacks bounds. Type is 'number' with no min/max constraints. LLMs may pass 0, negative, or extremely large values (e.g., 100000), causing performance issues or OOM errors.
search tool: source_type parameter is free-form string instead of enum. Description says 'Filter by file type (md/txt/pdf/xlsx)' but schema doesn't enforce these values. LLMs may pass invalid types like 'jpg' or 'doc', causing silent filtering failures or errors.
search tool: filters parameter object has no description. LLMs cannot determine when to use it, whether both sub-parameters are optional, or what happens if filters is omitted. The nested properties lack context.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 61 | <=2025-11-25 | v2 |
| 2026-03-09 | C | 62 | 2024-11-05+ | v1 |
No output schema documentation. Tool descriptions do not specify what fields search(), get(), or get_project_info() return. The code shows ToolResult with 'content' (array of ToolResultContent) and 'isError', but agents need to know: Does search() return ranked results with scores? What fields are in each result? What does get() return for a chunk, full text, metadata, embedding? What does get_project_info() contain?
search tool: No pagination support visible. Description does not mention limits on result count or pagination. If search returns hundreds of results, they are likely all returned as text, wasting tokens and LLM reasoning capacity. Should cap results at 20-50 and offer limit/offset parameters.
Error handling lacks recovery guidance. The code supports ToolResult.error() but tool descriptions do not explain what errors are possible or how LLMs should respond. E.g., 'chunk_id not found', should the agent search again? List available chunks? No guidance.