A read-only MCP server for Agentic RAG over Paperless-ngx documents. Enables semantic search and metadata filtering of document archives with LLM-powered fuzzy matching.
The server has 2 tools with reasonable structure but significant issues. get_current_date has a clear, focused purpose but returns plain text. search_paperless_metadata has extensive documentation in the description (addressing multiple failure modes and query syntax rules), but the schema is incomplete for what the description promises. Parameter schemas are present but lack type specificity for several fields (tags accepts string but description says 'comma separated tag IDs' without validation). The descriptions are lengthy (good for LLM clarity) but some encode implementation rules that should be constraints. Output format is not documented, both tools return strings rather than structured JSON. Error handling guidance exists in prose but not as formal error response types.
Returns today's date. Use this for precise date math (specific months, quarters, custom ranges). For common relative expressions like 'last year' or 'this year', you can instead pass time_range="last year" directly to get_paperless_master_data — no separate call needed.
Perform an exact keyword and metadata search in Paperless-ngx. Results are sorted NEWEST FIRST. The tool output always includes TODAY'S DATE so you can correctly resolve relative time expressions. IMPORTANT RULES: USE THIS TOOL when you know the EXACT Correspondent, Tag, or Document Type ID. To list the LATEST documents, leave the 'query' parameter empty. If the snippet or Custom Fields do NOT contain the needed detail, call `get_document_details` on those Document IDs to read the full OCR text. Do not give up early. Date parameters must be YYYY-MM-DD. Always call `get_paperless_master_data` first to resolve names to integer IDs. Do NOT pass string names to `correspondent`, `tags`, or `document_type` — only integer IDs. QUERY SYNTAX (CRITICAL): The Paperless full-text search uses AND logic: every word must appear in the document. NEVER pass multiple synonyms as one query. Use ONE precise keyword per call. For OR logic across terms, call this tool multiple times. Prefer `correspondent` or `tags` filters over a text `query` whenever possible. FALLBACK STRATEGY: If no documents found: retry with broader filters (e.g. remove date range). If still nothing: report which filters were applied and switch to `semantic_search_with_filters`. Never stay silent. Never ask the user for clarification before exhausting both search methods.
Output schema not documented. Both tools return plain strings rather than structured JSON. LLM cannot parse results without guessing at field structure.
Parameter type validation missing. The 'tags' parameter accepts a string but description says 'comma separated tag IDs', no validation enforces this format or rejects invalid IDs. The 'correspondent' and 'document_type' fields accept integers with description-only constraints but no min/max bounds.
Implementation rules encoded in description instead of constraints. The search_paperless_metadata description contains imperative rules ('Always call get_paperless_master_data first', 'Do NOT pass string names') that should be enforced via tool composition or parameter validation, not prose. This shifts burden to the LLM.
Inferred effective spec: 2026-07-28+.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 61 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 23 | - | v1 |
Missing tool for listing master data. The description references 'get_paperless_master_data' and 'semantic_search_with_filters' tools that do not appear in the registered tool list. If these are missing, the documented fallback strategies are unreachable.
No pagination limit enforcement in schema. page_size description says 'max 50' but schema does not enforce this with a maximum constraint. LLM can pass any integer, including absurd values.
Error handling guidance is prose-only. The description includes fallback strategies ('retry with broader filters', 'switch to semantic_search_with_filters') but these are not formalized as error response types or structured recovery hints the LLM can parse.