MCP server for Paperless-ngx REST API
Strong foundation with 21 well-named tools, comprehensive schemas, and clear descriptions. All tools follow verb_noun naming (paperless_get_*, paperless_list_*, paperless_search_*). Input schemas are complete with types and constraints (e.g., max_results capped 1-500, max_chars 100-200000). Descriptions are concise (50-150 chars typical) and action-oriented. Tool annotations (readOnlyHint, idempotentHint) present for all read tools. Main gaps: output schemas not documented in source; error handling guidance minimal; no confirmation pattern for the single write tool (paperless_acknowledge_task); parameter descriptions could be more prescriptive about format/constraints.
Mark a task as acknowledged (dismiss from UI).
Search documents and return excerpts with context. Useful for RAG/QA workflows.
Download a document as PDF/original. Prefer save_to_path for large files.
Fetch full document record (no OCR content).
Fetch OCR/extracted text for a document, plus key metadata. Use as RAG source.
Get audit history for a document.
Output schemas not documented in source code. LLMs cannot plan downstream tool calls or extract required fields without knowing response structure.
Taxonomy list tools (correspondents, types, storage_paths, tags) have minimal descriptions (10-20 chars). Descriptions should explain when to call each and what structure they reveal.
Single write tool (paperless_acknowledge_task) lacks confirmation/dry-run pattern. No guidance on error recovery or what happens on failure.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | A | 84 | 2026-07-28+ | v2 |
Get extended file metadata (mime, size, checksums, page count, lang). By default omits the verbose `original_metadata` / `archive_metadata` arrays. Set `include_raw_metadata=True` to receive them.
List notes attached to a document.
Get classifier suggestions (correspondent, tags, type, dates) for a document.
Get document thumbnail (typically small PNG). Prefer save_to_path for oversized blobs.
Get the next available archive serial number.
Get document statistics (counts by correspondent, type, tag, etc.).
Get details of a specific background task.
List all correspondents (senders/organizations).
List all document types.
List documents with structured filters. Auto-paginates up to max_results.
List all storage paths.
List all tags.
List background tasks (indexing, OCR, etc.).
Full-text search documents. Returns ranked hits with OCR highlights.
Verify Paperless credentials by fetching the current user profile. Returns authenticated user info on success or a structured error with the HTTP status on failure. Upstream response bodies are not echoed to the client; check server logs for full error details.
Parameter descriptions lack prescriptive format guidance. E.g., 'created_after' says 'ISO date YYYY-MM-DD' but doesn't validate or guide LLM on what happens with invalid dates.
Error handling responses documented in descriptions (e.g., 'check server logs for full error details') but no structured error classification (retryable vs user-fixable vs fatal).