A Model Context Protocol server for managing PDF documents with vector search capabilities, semantic embeddings, and hybrid search functionality
The server implements 11 tools with reasonable naming conventions and descriptions, but has significant gaps in schema completeness, parameter documentation, and error handling guidance. Tool names follow verb_noun patterns (list_documents, search_documents, add_document, etc.) which is good. Descriptions are present for all tools and range from 80-150 characters, meeting minimum length requirements. However, input schemas lack completeness: while parameter names and basic types are present, many parameters lack detailed descriptions of constraints, formats, and valid ranges. No output schemas are documented. Error handling is minimal, tools provide no recovery guidance or actionable error messages. The server lacks tool annotations (readOnlyHint, destructiveHint, idempotentHint) despite having clear risk levels assigned. Security considerations exist (non-root Docker user, environment variable configuration) but are not enforced in tool definitions via permission gates or scope declarations.
Add a new PDF document to the knowledge base from file path
Remove a document from the knowledge base
Export the knowledge base metadata and document index
Get detailed information about a specific document
Get or generate a summary for a document using the summarizer service
Get comprehensive statistics about the knowledge base
Import a previously exported knowledge base
List all documents in the knowledge base with optional filtering and pagination
No output schemas documented for any tool. LLMs cannot plan downstream calls or extract required fields (e.g., document_id from search results to pass to get_document_summary). Adds uncertainty and increases errors.
Missing tool annotations despite clear risk levels. delete_document (DESTRUCTIVE), add_document (WRITE), and import_knowledge_base (WRITE) lack readOnlyHint=false and destructiveHint flags. Agents cannot distinguish read-safe vs irreversible operations.
delete_document lacks confirmation/dry-run mechanism for destructive operations. An agent mistake could irreversibly delete documents from the knowledge base without recovery guidance.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 54 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 25 | - | v1 |
Scan the knowledge base directory for new or modified PDF files and update the index
Search documents using vector similarity and optional full-text search with reranking
Update a document in the knowledge base by replacing it with a new version
Parameter descriptions lack constraint details. 'sort_by' accepts string with no enum of valid values (name, date_added, size), forces LLM to guess or invent values. 'limit' parameters have no min/max bounds. 'query' in search_documents has no guidance on format or max length.
No error handling guidance. Tools provide no recovery instructions. E.g., if add_document fails with 'file not found', there's no message like 'Verify file_path is absolute and readable.' Error responses cannot guide agent recovery.
No pagination guidance for list_documents. Tool accepts skip/limit but doesn't document total count return, next_cursor, or max result limits. Large knowledge bases could return unbounded lists, exhausting context windows.
Parameter 'file_path' in add_document and update_document accepts arbitrary strings with no validation guidance. No mention of path traversal safety, absolute vs relative paths, or supported file systems. Agents could pass malicious paths.
Tool composition risk: get_document_summary accepts document_id directly, but search_documents likely returns document references that may not be named 'document_id'. Undocumented field mismatch forces agent to infer ID field names.