HTTP API and MCP server for converting documents, web pages, and media to markdown
Single tool 'convert_to_markdown' has a well-structured schema with comprehensive parameter coverage and clear descriptions. The tool definition is explicitly visible in src/md_server/mcp/tools.py and properly registered via FastMCP decorator in server.py. Naming is action-verb-based and descriptive. Description is substantial (~280 chars) and explains use cases, supported formats, and output modes. Parameter schema is complete with types, descriptions, enums, and sensible defaults. Output format is documented (markdown or json). However, output schema structure itself is not formally documented, and error handling lacks specific recovery guidance beyond what the implementation provides.
Read a URL or file and convert to markdown. Provide ONE of: - url: Webpage, online PDF, or Google Doc - file_content + filename: Base64-encoded file Supported formats: PDF, DOCX, XLSX, PPTX, HTML, images (OCR), and more. For JavaScript-heavy pages (SPAs, dashboards), set render_js: true. This adds ~15-30 seconds but captures dynamically loaded content. Returns markdown by default, or structured JSON with metadata (set output_format: "json").
Output schema not formally documented in tool definition. While output_format enum declares 'markdown' or 'json', the structure of returned markdown or JSON objects is not formally specified. LLMs cannot plan downstream processing without knowing what fields/structure to expect.
Error recovery guidance incomplete. Tool description does not explain what errors can occur (invalid URL, unsupported format, timeout, auth failure) or how to recover. LLM receives raw error but no actionable next steps.
No idempotency guarantee documented. Tool reads external URLs and files, repeated calls with same input may produce different results (page content changes, temp file deleted). Agents retry on ambiguous failures; without idempotency guarantees, this risks duplicate or inconsistent behavior.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-21 | C | 64 | 2026-07-28+ | v2 |
| 2026-03-09 | C | 64 | - | v1 |
Parameter truncate_limit lacks min/max constraints. Description says it's a 'limit for truncation mode' but does not bound the value. An agent could pass truncate_limit=999999999, causing memory/performance issues.
Mutually exclusive parameters (url vs file_content+filename) not formally documented as such. Description says 'Provide ONE of', which is good, but parameter schema does not enforce this constraint. LLM may pass both, causing undefined behavior.