Convert PDF documents to clean Markdown with text, tables, images, and equation support. Exposes PDF-to-Markdown converter as MCP tools, resources, and prompts for integration with VS Code GitHub Copilot, Claude Desktop, Cursor, or any MCP-compatible AI client.
The pdf2md MCP server defines 2 tools with complete, well-structured input schemas and detailed parameter descriptions. Both tools have clear, substantive descriptions (150+ chars each) that explain what they do and when to use them. However, there are notable gaps: (1) output schemas are defined via Pydantic BaseModel but not documented in tool descriptions; (2) error handling is implicit (rely on exceptions) with no recovery guidance in descriptions; (3) no tool annotations (readOnlyHint/destructiveHint/idempotentHint); (4) sensitive parameters like 'api_key' are exposed as input params rather than server-side injected; (5) enum constraints are documented in descriptions but not formally declared in JSON Schema. Per-tool analysis: convert_pdf (score 68) has 23 input parameters with full type definitions and descriptions, but parameter relationships (e.g., llm_mode='none' makes api_key irrelevant) are not explicitly documented. convert_pdf_batch (score 68) mirrors convert_pdf with identical parameter structure and the same gaps. Both tools follow verb_noun naming (convert_pdf, convert_pdf_batch) which is appropriate.
Convert a single PDF file to Markdown. Extracts text with heading detection, tables, embedded images, and renders math equations as images automatically. Optionally use a vision LLM for enhanced extraction.
Convert all PDF files in a directory to Markdown. Scans the given directory for PDF files matching the pattern and converts each one to a .md file in the output directory.
Sensitive parameter 'api_key' exposed as tool input. This violates secret-injection pattern, credentials must never appear as parameters since agent traces log all parameter values.
Enum constraints (llm_mode, image_llm_mode, equation_llm_mode, image_detail_level, engine) are documented in descriptions as free-form strings, not declared as JSON Schema enums. LLMs cannot reliably parse prose constraints and may hallucinate invalid values.
No error handling guidance. Descriptions do not explain what exceptions might be raised (e.g., file not found, invalid password, LLM API failure) or how the LLM should recover. Pattern: recovery-guide expects error responses to tell the agent what to do next.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 69 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 31 | - | v1 |
Output schema (SingleConversionResult, BatchConversionResult) is defined via Pydantic models but not documented in tool descriptions. LLMs cannot infer that convert_pdf returns markdown, output_path, images_dir, etc. without this information in the description.
Parameter relationships are undocumented. E.g., llm_mode='none' makes api_key, api_base_url, text_model, vision_model, etc. irrelevant, but this dependency is not declared.
No tool annotations. Tools do not declare destructiveHint for state-modifying operations (write to disk) or idempotentHint. This prevents MCP clients from applying safety policies or caching.
Numeric parameters lack explicit min/max constraints. E.g., render_dpi, min_image_size, max_tokens, temperature have no documented bounds. LLMs may pass absurd values (e.g., render_dpi=1000000, temperature=999) that break logic or APIs.