Static source inference · medium confidence · evidence: Streamable HTTP
Current-spec patterns detected
Summary
PaperlessMCP exhibits strong tool naming conventions and comprehensive schema definitions across 29 tools. All tools use verb_noun patterns (list, get, create, update, delete, search) that clearly convey intent. Input schemas are fully visible and typed. However, descriptions vary significantly in quality: some are concise and actionable (10-30 chars), while others lack specificity about when to use them or what prerequisites exist. Error handling is present but inconsistent, some tools (delete operations) include confirmation patterns, but others lack recovery guidance. Output schemas are not explicitly documented in the source, making it difficult for LLMs to understand response structure beyond inference from API responses. Tool composition is sound, each tool has a single responsibility, and references are preserved across tool chains (e.g., correspondent_id in documents).
Output schemas are not explicitly documented in source code. While tools return structured responses via McpResponse<T> wrappers, the actual field schemas for each response type (Correspondent, DocumentType, etc.) are not visible in the provided code. LLMs must infer response structure from parameter context rather than explicit documentation.
Generic descriptions on GET and simple read operations lack context about when to use them vs. alternatives. E.g., 'Get a correspondent by its ID' does not explain whether to use this after searching or for single-ID lookups. Descriptions should include when to call this tool relative to similar operations.
Expose response schemas explicitly. Create a schema document or inline JSON Schema definitions showing the structure of Correspondent, Document, CustomField, etc. Include field types, nested object schemas, and whether fields are optional. This is critical for LLMs to understand what data they receive and plan multi-step chains.
Enhance GET tool descriptions to include when to use them relative to list/search operations. E.g., 'Get a correspondent by its ID. Use after searching for correspondents with paperless_correspondents_list, or if you already know the correspondent ID.'
Add recovery guidance to non-destructive error responses. When creation or update fails, include suggestions: 'Correspondent creation failed: name already exists. Use paperless_correspondents_search to find the existing correspondent, or use a different name.'
Document actual MAX_PAGE_SIZE and result caps in pagination descriptions. State: 'Page size is capped at 100. Requesting larger sizes will be silently reduced.'
Add dependency hints to discovery tools. E.g., 'List all correspondents. Call this first to see available correspondents before assigning them to documents, or to find a correspondent ID for filtering documents.'
Get download URLs for a document's original file, preview, and thumbnail. Set returnBase64=true to also inline the file bytes as base64, but only for tiny files: use paperless_documents_export_to_outbox for anything real.
Error handling does not consistently provide recovery guidance. While destructive operations (delete, bulk_delete) include confirmation patterns and return structured error responses, read and write operations (create, update) lack actionable next-step guidance. An LLM encountering a 404 or validation error has no clear recovery path suggested.
Matching algorithm parameters use integer enums (0=None, 1=Any, 2=All, 3=Literal, 4=Regex, 5=Fuzzy, 6=Auto) without exposing them as a constrained enum in JSON Schema. This forces LLMs to remember the mapping or guess values. Should declare as enum: [0, 1, 2, 3, 4, 5, 6] with human-readable labels in the description or use string enum values.
Custom field 'dataType' parameter lists valid values in description ('string, url, date, boolean, integer, float, monetary, documentlink, select') but does not expose them as a JSON Schema enum. This invites hallucinated values and requires LLM to parse the description string.
Pagination support exists (page, pageSize, ordering parameters) but result limits are not enforced or documented. The description mentions 'capped by MAX_PAGE_SIZE' but does not state the actual limit. Without documented caps, agents may request pages that exceed reasonable bounds.
Verify that bulk_delete operations with comma-separated IDs validate the input format and return clear errors if parsing fails (e.g., 'Invalid ID list: expected comma-separated integers, got: "1,2,invalid"'). Document this validation behavior.
For custom field assignment (paperless_custom_fields_assign), document the format of the 'value' parameter more explicitly: 'Value to assign. Format depends on fieldType: string for string/url/date fields, boolean (true/false) for boolean fields, comma-separated document IDs for documentlink fields. Example: For a date field, pass "2024-01-15"; for documentlink, pass "123,456,789".'
Add tool annotations (readOnlyHint, destructiveHint) to leverage protocol capabilities. Mark all read-only tools with readOnlyHint=true; mark all delete/bulk_delete tools with destructiveHint=true. This helps clients and LLMs understand operation safety.