Citation-ready PDF and DOCX workflows that export reusable agent assets and Foam/LightRAG wikis over MCP
Asset-Aware MCP has 4 tools with partial schema and description coverage. Two tools (ingest_docx, get_docx_content) have detailed descriptions in Chinese/English mix, but schemas are minimally specified, parameter types are present but lack comprehensive validation constraints. The other two tools (get_job_status, list_jobs) have generic, minimal descriptions. No tools declare output schemas or error handling guidance. Parameter descriptions are sparse; input constraints (min/max, enums) are not enforced. No tool declares permissions, idempotency guarantees, or whether calls are retryable. The server's STDIO-only transport limits remote accessibility, but definition quality issues persist independently.
取得 docx 文件的可編輯 DFM 內容。 若指定 block_id,只回傳該區塊的內容;否則回傳完整 DFM。
Get the status of an ETL job. Use this to check progress of document ingestion started with `ingest_documents`.
攝入 .docx / .doc 文件,轉換為 DFM (Docx-Flavored Markdown) 格式。 將 docx 解析為中間表示 (IR),再轉換為可在 VS Code 中編輯的 DFM 格式。 支援複雜元素:合併表格、圖表、頁首頁尾、巨集、目錄等。 **支援舊版 .doc 格式**(自動透過 LibreOffice 轉換為 .docx)。 輸出目錄結構: ``` data/{doc_id}/ ├── content.dfm # 可編輯的 Markdown + YAML 標注 ├── ir.json # IR 快照(用於回寫) ├── original.docx # 原始檔案備份 ├── parts/ # 保留的 XML 零件 └── assets/ # 圖片和二進位資產 ```
List ETL jobs.
Output schemas not documented. LLMs cannot plan downstream tool calls without knowing return structure. ingest_docx, get_docx_content, get_job_status, and list_jobs all lack explicit return type documentation.
Minimal parameter descriptions. get_job_status has only 'Job ID returned from `ingest_documents`' (bare, no format hints). list_jobs has only 'If True, only show pending/processing jobs' (minimal context for when to use). LLMs cannot disambiguate parameter semantics.
No error handling guidance. None of the 4 tools specify what errors are retryable, user-fixable, or fatal. No recovery suggestions (e.g. 'file not found, check the file_path parameter'). Agents cannot self-correct on failure.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 46 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 45 | - | v1 |
No input constraints or validation hints. ingest_docx accepts file_path as string with no format validation (absolute path? relative? glob patterns?). get_docx_content's max_chars parameter has no min/max bounds. Unbounded parameters allow agents to pass absurd values.
Incomplete or vague descriptions undermine discoverability. get_job_status (25 chars) and list_jobs (22 chars) fall well below the 50-200 char baseline for LLM-optimized descriptions. They do not explain WHEN to call these tools or WHAT they return.
No idempotency or side-effect declarations. ingest_docx is marked WRITE but lacks clarity on whether repeated calls with the same file_path re-process or return cached results. No confirmation/dry-run pattern for destructive operations.
Tool composition incomplete. ingest_docx outputs doc_id, but get_docx_content's block_id parameter description does not explain the format (e.g., are valid block_ids only 'p001', 't001', 'h001', or can agents invent them?). Ambiguity breaks tool chaining.
Pagination not implemented. list_jobs has no offset/limit/cursor parameters documented, raising questions about scalability and context bloat if the result set is large.
No permission/scope declarations. Tools do not advertise required permissions (e.g., 'read:docx', 'write:data'). Cannot verify least-privilege agent configuration.
Mixed-language descriptions (Chinese + English) may confuse LLMs or clients expecting ASCII descriptions. Consistency is key for multi-model deployments.