Local CLI and MCP server for PDF form filling, editing, OCR, and AI-assisted extraction
The pdf-mcp server provides 34 tools with consistent naming patterns (verb-noun: get_, create_, extract_, etc.) and reasonable descriptions. However, there are critical gaps: most tools lack detailed input parameter descriptions in the visible code, output schemas are not documented in the submission, and error handling guidance is absent. The server implements core PDF operations competently but falls short of production-grade agent tooling standards. Tool definitions appear in registry.py and pdf_tools.py with schemas visible in the submission, but descriptions are often single-line summaries (e.g., 'Return available form fields in the PDF') that lack guidance on when to use each tool vs. similar ones, prerequisites, or recovery paths. Parameter schemas are defined (all tools have 'type' and 'description' fields in input objects), but descriptions are minimal (e.g., 'Path to the PDF file' repeated 60+ times without nuance). Output schemas are not visible in the submission and are not documented in the code samples provided.
Add page numbers to a PDF.
Add a digital signature to a PDF.
Add a FreeText annotation to a page (managed text insertion).
Add a watermark text to all or specific pages.
Clear (delete) values for PDF form fields while keeping fields fillable.
Compare two PDFs and report differences.
Create a new PDF with AcroForm fields.
Output schemas are not documented for any tool. The submission does not include return type definitions, field lists, or examples of what each tool returns to the LLM. This violates pattern:tool requirement that all tools must document outputs so agents can parse responses and plan downstream actions.
Parameter descriptions are generic and repetitive (e.g., 'Path to the PDF file' used identically across 30+ tools). There is no guidance on when to use fill_pdf_form vs fill_pdf_form_any, what makes a PDF 'non-standard', or how label detection differs from AcroForm field mapping. LLMs cannot distinguish between similar tools without explicit comparative context.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 57 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 34 | - | v1 |
Create a PDF form using a built-in template.
Detect PDF type and features (AcroForm, XFA, encrypted, etc.).
Encrypt (password-protect) a PDF using pypdf.
Export PDF content and metadata to JSON format.
Export PDF content to Markdown format.
Extract images from a PDF.
Extract hyperlinks and their target URLs from a PDF.
Extract specific 1-based pages into a new PDF.
Extract tables from a PDF.
Extract text and their bounding boxes from a PDF.
Fill a PDF form with provided data. Optionally flatten to make non-editable.
Fill standard or non-standard forms using label detection when needed.
Flatten a PDF (remove form fields/annotations).
List built-in form templates for common workflows.
Return available form fields in the PDF.
Extract and return PDF metadata (title, author, creation date, etc.).
Use an LLM to analyze a PDF document and answer questions about it.
Use an LLM to extract structured data from a PDF.
Use an LLM to auto-fill a PDF form based on context or document content.
Merge multiple PDFs into a single file.
Extract text from a PDF using OCR (Tesseract).
Redact (black out) specified text from a PDF.
Reorder pages in a PDF using a 1-based page list.
Rotate specified 1-based pages by degrees (must be multiple of 90).
Scan a PDF for personally identifiable information (PII) patterns.
Update PDF metadata (title, author, subject, etc.).
Verify digital signatures in a PDF.
Tool descriptions lack guidance on LLM tool selection. Descriptions are single sentences (e.g., 'Flatten a PDF (remove form fields/annotations)') without explaining when flatten_pdf is necessary, whether it is reversible, or what happens to interactivity. Descriptions should answer: What does it do? When to use it? What are the side effects?
No error handling guidance. Tools like add_signature, encrypt_pdf, and redact_text (IRREVERSIBLE operations) do not document recovery paths or error conditions. An LLM receiving 'cryptography.hazmat.bindings.openssl.binding._OpenSSLErrorWithText' has no guidance on whether to retry, ask the user, or abort.
LLM-integration tools (llm_fill_form, llm_extract_data, llm_analyze_document) do not document which LLM provider is used, how to inject credentials, or whether they are stateless. The description 'Use an LLM to auto-fill a PDF form based on context...' hides complexity about OpenAI API key handling, cost, latency, and fallback behavior when the LLM service is unavailable.
No pagination or result limits documented. Tools like extract_tables, extract_images, and ocr_pdf can potentially return very large results (hundreds of table objects or thousands of image references). Without documented limits (e.g., 'Returns max 100 tables per call; use page-range to limit scope'), LLMs may inadvertently request 500-page PDFs, exhausting context windows.
Tool naming ambiguity: fill_pdf_form vs fill_pdf_form_any. The naming does not indicate that the second variant auto-detects non-AcroForm fields. Per pattern:tool-naming, similar tools must have clearly distinct names (e.g., fill_pdf_form_acroform vs fill_pdf_form_heuristic) or consolidated into one smart tool with a parameter controlling detection mode.
Credentials for signature operations (key_path, cert_path in add_signature) are exposed as tool parameters. Per pattern:secret-injection, cryptographic keys should never be passed as parameters, they should be injected server-side via environment variables or a key vault. Passing them as params risks leaking secrets into agent trace logs.
No scope declarations or permission gates. Tools that modify sensitive documents (encrypt_pdf, redact_text, add_signature, flatten_pdf) do not declare required permissions or check caller authorization. An agent should not be able to redact text from a confidential document without explicit permission.
Parameter 'fields' in clear_pdf_form_fields is under-described. Description says 'Optional list of specific field names to clear; if omitted, all fields are cleared' but does not explain: (1) What happens if a field name does not exist? (2) Are field names case-sensitive? (3) How are nested field hierarchies referenced (e.g., 'parent.child')? (4) What is returned on success, the list of cleared fields?
Parameter 'schema' in llm_extract_data lacks format constraints. Description says 'JSON schema for the data to extract' but does not specify: (1) Does it expect JSON Schema Draft 7 or OpenAPI 3.0 format? (2) Are custom formats (e.g., 'date-time', 'email') supported? (3) What happens if the schema is invalid or unsatisfiable? This forces the LLM to guess at the format.