A Model Context Protocol (MCP) server that enables LLMs (like Claude or any other LLM that supports MCP) to query your documents using semantic search. Organizes documents by topics based on folder structure. Supports: PDF, Word, Excel, Markdown.
The server has 8 well-named tools with reasonable descriptions and documented input schemas. Tool names follow verb_noun conventions (search_documents, list_documents, get_collection_stats). Most parameters have type definitions and descriptions. However, output schemas are not explicitly documented in the source code, only input schemas are visible. Error handling guidance is absent. Several tools lack detail on return structure, field mapping, and recovery paths. The schema quality is moderate but not exceptional. Descriptions are adequate (100-150 chars average) but could be more specific about when to use each tool and what downstream tools they enable.
Get statistics about the document collection.
Get the timestamp of when the last folder scan was started and finished.
Get a list of all available documents with their hierarchical topics.
Get a list of all topics/categories in the document collection.
Scan all documents in the docs directory and update the vector database.
Search across all documents using semantic similarity. Returns the most relevant text chunks with their source information and hierarchical topics. Optionally filter by topic, date range, or regex pattern.
Output schemas not documented. Tools return strings (based on code analysis of mcp_tools.py), but the response structure, field names, and chaining IDs are not specified. LLMs cannot plan downstream calls or extract structured data.
No error handling guidance. Tools lack descriptions for failure cases, recovery steps, or what actions to take if a scan fails, a search returns no results, or folder watching cannot start. Bare error messages will not guide LLM recovery.
Incomplete parameter descriptions for filter parameters. 'date_from' and 'date_to' lack format specification (Unix timestamp vs ISO 8601). 'regex_pattern' does not specify regex syntax (Python re vs PCRE). LLMs will guess and pass malformed input.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 66 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 12 | - | v1 |
Start watching the documents folder for changes and automatically trigger incremental updates.
Stop watching the folder for changes.
No output pagination or result limits explicitly stated for search_documents. While max_results parameter exists, the description does not clarify if results are ranked by relevance, how many fields each result contains, or whether the response includes a total_count.
Watched folder state management is stateful but undocumented. Tools start_watching_folder and stop_watching_folder have side effects that persist across calls, but descriptions don't clarify initial state, idempotency, or what happens if stop is called when not watching.
Tool discovery incomplete. Descriptions don't explain when to call list_documents vs list_topics, or which should be called first. No guidance on tool orchestration or when scan_all_my_documents should be triggered vs start_watching_folder.