MCP server for local document indexing and search using LanceDB
The server provides 4 tools with explicit schema definitions (via Pydantic models) and descriptions. Naming follows verb_noun convention (search_documents, get_catalog, get_document_info, reindex_document). All tools have descriptions ranging from 180-350 characters, within the 10-1024 character target. Input parameters have types and descriptions. However, several quality gaps prevent a higher score: (1) output schemas are documented only informally in code comments, not as structured response definitions; (2) descriptions for parameters like 'file_path' in get_document_info lack format guidance; (3) error handling is generic ('success: false, error: str(e)') without actionable recovery guidance per the pattern:recovery-guide; (4) no error categorization (retryable vs fatal); (5) no tool annotations (readOnlyHint, destructiveHint, idempotentHint) despite clear risk differences (3 READ_ONLY, 1 WRITE); (6) sample queries in search_documents descriptions are helpful but examples in parameters should be avoided. The Pydantic models are well-structured and provide solid input validation. Output structures are reasonable (returning success, total_results, formatted results) but lack formal schema documentation for downstream chaining.
Get a list of all indexed documents with their summaries. Returns a catalog of all documents that have been indexed, including their metadata, summaries, and keywords. Useful for browsing the document collection. Sample usage: - Browse all indexed documents to see what's available - Get an overview of the document collection with metadata - Check which file types are indexed and their summaries - Find documents by scanning through titles and keywords
Get detailed information about a specific indexed document. Returns comprehensive information about a document including its summary, keywords, chunk count, and indexing metadata. Sample usage: - Get details about "/path/to/my/important-document.pdf" - Check metadata for "src/main.py" including chunk count and size - View summary and keywords for "docs/README.md" - Inspect indexing status and timestamps for any file
Reindex a document to update its embeddings and metadata. Forces reindexing of a specific document even if it appears unchanged. Useful when you want to refresh embeddings or apply new processing logic. Sample usage: - Reindex a specific document to update its embeddings - Force refresh of summary and keywords for a document - Apply new processing to an existing document without deleting it
Search for documents or chunks using semantic search. This tool searches through indexed documents using natural language queries. It can search at the document level (returning whole documents) or chunk level (returning specific passages). Sample queries: - "Find documents about machine learning algorithms" - "Search for API documentation" - "Show me code related to database connections" - "Find text about authentication and security" - "Look for configuration files and setup instructions"
Output schemas not formally documented. Tool responses include structured objects (results arrays, metadata, stats) but the schema of returned fields is not declared in tool definitions. LLMs cannot infer downstream parameter passing without explicit response schemas.
Error handling lacks actionable recovery guidance. All tools return {'success': false, 'error': str(e)} on failure. Errors do not indicate whether to retry, ask the user, or abandon. Per pattern:recovery-guide, errors should guide the next step.
Tool annotations missing. Three tools are READ_ONLY (search_documents, get_catalog, get_document_info) and one is WRITE (reindex_document), but tool definitions do not include readOnlyHint or destructiveHint annotations. Agents cannot distinguish safe reads from mutating operations without these hints.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 68 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 47 | - | v1 |
Parameter 'file_path' lacks format guidance. get_document_info and reindex_document accept 'file_path' but do not specify format (absolute? relative? pattern? max length?). Per pattern:constrained-input, format constraints prevent hallucinated values.
get_catalog and reindex_document lack pagination context. reindex_document takes a single file_path and has no batch variant. get_catalog accepts skip/limit but description does not state the maximum result count or warning about large result sets. Per pattern:paginated-result and mxe:enforce-result-limits, paginated tools should declare limits.
search_documents includes example queries in description ('Find documents about machine learning algorithms'). Per pattern:tool-description, examples bias LLMs to reuse them literally rather than adapting to context. Move examples to separate documentation or replace with enum constraints.