Enterprise AI Data Preprocessing Platform - MCP Server for document, image, audio, and video processing with RESTful API and Model Context Protocol interfaces
MinerU Tianshu exposes 4 tools via HTTP MCP with reasonable structure. Tool naming follows verb_noun conventions (parse_document, get_task_status, list_tasks, get_queue_stats). All tools have descriptions in Chinese (complete, 100-300 chars), and input schemas are present with typed parameters. However, descriptions lack English equivalents and explicit guidance on when/why to use each tool. Parameter descriptions are present but terse. Output schemas are not documented in the provided source, making it impossible to verify what fields agents should expect. No error handling guidance is visible. The server supports sensible features (priority, async processing, status filtering) but lacks annotations (readOnlyHint/destructiveHint) and comprehensive error recovery patterns. Schema completeness is moderate, all parameters are typed with enums where appropriate, but output structure is opaque.
获取任务队列统计信息。 返回各个状态的任务数量,了解系统负载情况。
查询文档解析任务的状态和结果。 可以查询任务的: - 当前状态(pending/processing/completed/failed/cancelled) - 处理进度和时间信息 - 错误信息(如果失败) - 解析结果内容(如果完成)
列出最近的文档解析任务。 可以按状态筛选,查看任务队列情况。
解析文档(PDF、图片、Office文档等)为 Markdown 格式。 📁 支持 2 种文件输入方式: 1. file_base64: Base64 编码的文件内容(推荐用于小文件) 2. file_url: 公网可访问的文件 URL(服务器会自动下载) 支持的文件格式: - PDF 和图片(使用 MinerU GPU 加速解析) - Office 文档(Word、Excel、PowerPoint) - 网页和文本(HTML、Markdown、TXT、CSV) 功能特性: - 公式识别和表格识别 - 支持中英文、日文、韩文等多语言 - 支持任务优先级设置 - 异步处理,可选择等待完成或稍后查询
Output schemas are completely undocumented for all 4 tools. Agents cannot plan downstream calls or extract required data (task IDs, statuses, counts, markdown content). This forces LLMs to guess at response structure and risks parsing errors.
All descriptions are in Chinese only. Multi-language tool descriptions are impractical; English descriptions are standard for LLM-driven discovery and error handling. Non-English descriptions reduce tool discoverability and force LLMs to translate on the fly (error-prone).
Error handling and recovery guidance is absent. No documentation of failure modes (invalid file format, timeout, task not found, queue full, permission denied). Tools return errors but agents lack context to retry, escalate, or self-correct. E.g., parse_document silently failing on unsupported format with no guidance.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 67 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Tool annotations (readOnlyHint, destructiveHint, idempotentHint) are absent. Agents cannot infer which operations are safe to retry. parse_document with wait_for_completion=true is idempotent (same file, same params = same result); this is not declared, causing agents to over-retry or under-retry.
list_tasks output pagination strategy is undocumented. If results exceed 100 items, agents cannot know if truncation occurs, whether next_cursor is provided, or how to fetch subsequent pages. This breaks composition for monitoring large task queues.
parse_document supports 8 backends but lacks clear guidance on when to use each (pipeline vs vlm-auto-engine vs hybrid-auto-engine, etc.). LLMs will default to 'pipeline' without understanding tradeoffs. Backends should have brief descriptions in enum.
Parameter 'max_wait_seconds' (10-3600s) has no guidance on timeout behavior. Does the tool return partial results? Error with task_id for async retry? Hang and time out? Agents cannot tune this intelligently.