MCP server providing document processing tools including image extraction from DOCX files, text extraction, table extraction, ZIP asset extraction, image tagging with OCR, local image analysis, and thesis reference management
This server has 14 tools with significant quality gaps. Most tools lack adequate descriptions (many under 50 chars), parameter schemas are present but descriptions are sparse or generic, and error handling is minimal. Tool naming is generally verb-forward (extract_*, tag_*, search_*, analyze_*) which is good, but several tools have overly terse or non-English descriptions. The server mixes document processing (docx, zip), image analysis, and academic paper management, disparate domains with weak cohesion. Many tools require 'user_working_dir' as a mandatory parameter, which is documented in comments but not enforced at schema level. Output schemas are largely undocumented, callers must infer structure from implementation. Error handling returns basic dicts with 'error' keys but provides minimal recovery guidance.
分析论文中使用的引用,找出未使用的参考文献
分析单个图像并生成标题
批量分析目录中的图像
删除未使用的参考文献
将LaTeX中的\cite引用转换为上标格式
Extract images from a .docx file into output_dir. Return list and meta.
Extract tables and embedded Excel content from .docx file with structured data format.
Six tools (analyze_single_image, batch_analyze_images, search_papers_arxiv, get_paper_details, analyze_citations, clean_unused_references) have descriptions under 50 characters, violating the 10-1024 character guideline. Chinese-language descriptions ('分析单个图像并生成标题', '批量分析目录中的图像', etc.) provide no context in English-speaking agent environments.
Parameter descriptions are missing or generic. For example, 'analyze_single_image' has parameters 'image_path', 'user_working_dir', 'context' with descriptions like '图像文件路径' (image file path) and '上下文信息' (context information), no guidance on format, required length, or role. 'extract_docx_images' parameter 'user_working_dir' is described as 'User's working directory (REQUIRED - ask user for this)' but this is a comment, not a constraint that prevents invalid paths.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 45 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 42 | - | v1 |
Extract text from .docx file with optional character limit. Now includes table and Excel content.
Extract assets from a ZIP file into output_dir.
获取论文详细信息
hello world
保存搜索结果到指定目录 - 按领域分类保存到 references/领域/ 目录
在arXiv上搜索论文
Tag exported images with OCR text and metadata.
Output schemas are entirely undocumented. For example, 'extract_docx_images' returns a dict with keys 'error', 'suggestion', 'base_directory', but the tool docstring does not describe the success case structure. 'extract_docx_tables' returns a list of dicts with 'table_index', 'rows', 'columns', 'data', but no schema is visible to callers. This forces LLMs to infer output structure from implementation details.
Error handling is minimal and unhelpful. Errors return dicts like {'error': 'File not found: ...', 'suggestion': '...', 'base_directory': '...'}, but do not categorize errors as retryable vs. fatal, and do not guide the LLM on recovery actions beyond vague suggestions. No tool provides an enum of possible error codes or structured error classifications.
The 'helloworld' tool is a stub ('hello world' description, no meaningful functionality) and pollutes the toolset. It should be removed.
Tool composition is poor. This server spans three unrelated domains: document processing (docx, zip extraction), local image analysis, and academic paper search/management. No coherent domain model. LLMs will struggle to discover which tool to call when working with papers vs. images vs. documents.
Input validation is weak. Many tools accept file paths but do not validate format (e.g., _validate_path in docx_image_tagger checks existence but does not prevent path traversal). The 'user_working_dir' pattern requires users/LLMs to pass their project directory every call, this is error-prone and not enforced by schema constraints.
Many parameters could benefit from enums or constraints. For example, 'ocr_lang' defaults to 'chi_sim+eng' but accepts free-form strings; 'source' in 'get_paper_details' has an enum ['arxiv', 'doi', 'title'] which is good, but most other tools lack enums. 'max_chars' in 'extract_docx_text' is an integer with no bounds (default 50000), an LLM could pass 1000000 and cause a timeout.