A modular, pluggable RAG (Retrieval-Augmented Generation) service framework with MCP protocol support for semantic search and document ingestion
The server has 5 tools with clear names and moderately detailed descriptions in Chinese. Tool descriptions are action-oriented and explain when to call each tool, this is good practice. However, there are significant gaps: parameter descriptions are minimal (mostly just type/optionality), input schemas lack detailed constraints, no documented output schemas are visible in the source, and error handling guidance is absent. The tools themselves are well-named (query_knowledge_hub, list_collections, ingest_document_*) following verb_noun convention. Parameter schemas are present in the function signatures but descriptions within those schemas are sparse. No enum constraints, validation rules, or recovery guidance appear in the code. Output structure is not formally documented. This places the server in the 'fair' range: core definitions exist but fall short of production quality.
当用户要把文档加入知识库、且文档包含较复杂表格、标题层级、图片混排,但仍是文本型文档时调用。使用 Docling 做结构化解析和 HybridChunker 切块,并写入向量库。参数: file_path(必填), collection_name(可选)。
当用户要把 PDF 加入知识库、或要求用 MinerU 精细解析/入库时调用。使用 MinerU 云端 API 解析 PDF 并写入向量库,适用于表格多、公式多、扫描件等复杂排版。参数: file_path(必填), collection_name(可选)。
当用户要把 PDF 加入知识库、且要求普通/快速解析(非 MinerU)时调用。使用本地 PdfLoader(MarkItDown)解析 PDF 并写入向量库,适用于简单排版 PDF,无需云端。参数: file_path(必填), collection_name(可选)。
当用户想知道有哪些知识库集合、或需要选择/确认 collection 时调用。列出已入库的集合名称。
当用户询问知识库内容、文档相关问题、或需要检索资料时调用。在 RAG 知识库中语义检索,返回相关文档片段(Markdown)和引用。参数: query(必填), collection_name(可选), top_k(可选,默认10)。
Parameter descriptions are minimal and lack constraint details. For example, 'top_k' is described only as '返回结果数量(可选,默认10)' with no bounds (min/max), no clarification that higher values increase latency/tokens, and no guidance on typical ranges.
No documented output schemas. The source code does not show what query_knowledge_hub returns (e.g., does it return document_id, snippet, confidence_score, source_file?). LLMs cannot plan downstream tool calls or extract required data without knowing output structure.
No error handling or recovery guidance in tool descriptions. What happens if file_path does not exist? If collection_name already exists? If the PDF is corrupted? Descriptions should include 'If file not found, use...' or similar recovery hints.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 48 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 30 | - | v1 |
list_collections takes no parameters and has no documented output. Does it return a flat list of strings? A structured array with metadata (size, created_date, document_count)? Without documented output, the LLM cannot extract IDs for downstream calls.
No enum constraints for file paths or collection names. 'file_path' is described as a string but has no validation hints (must be absolute? relative to cwd? must exist before calling?). This invites malformed input.
Descriptions use Chinese exclusively. While not a technical defect, this limits usability for non-Chinese-speaking agents and violates the convention that descriptions should be in English for broad compatibility with LLM frontends.