MCP智能知识库助手 - A Retrieval-Augmented Generation (RAG) system using Aliyun Qwen LLM and ChromaDB vector store for MCP protocol knowledge management
This RAG MCP server has significant definition quality gaps. While tool schemas are present with basic type information, descriptions are often translated from Chinese and lack actionable guidance for LLM selection. Most critically, output schemas are entirely undocumented, the server provides no specification of what fields agents should expect from tool responses. No parameter constraints (enums, ranges) are declared. Error handling lacks recovery guidance. The tool composition shows moderate coherence but overlapping concerns (build_knowledge_base, add_documents, and process_documents all manipulate the knowledge base with unclear responsibilities). No permission gates, audit trails, or security hardening visible despite destructive operations (clear_collection, build_knowledge_base with clear_existing=true). Tool names are reasonably action-oriented (build_, generate_, search_, add_, get_, clear_) but descriptions are sparse and sometimes ambiguous, e.g., 'Searches the knowledge base' doesn't explain when to use search vs generate_response.
添加文档 - Adds documents to the vector store with embeddings generated via Aliyun Qwen Embedding API
构建知识库 - Builds the MCP knowledge base from text documents in the txt directory, optionally clearing existing data and using header-based or size-based text splitting
清空集合 - Clears all documents from the ChromaDB collection and reinitializes it
生成回答 - Generates an answer to a user query using RAG by retrieving relevant documents from the knowledge base and generating a response via Aliyun Qwen LLM
获取集合信息 - Retrieves information about the ChromaDB collection including document count and storage path
获取嵌入向量 - Generates text embedding vectors using Aliyun Qwen Embedding model (text-embedding-v4)
Output schemas entirely undocumented across all 10 tools. LLMs cannot infer response structure, forcing agents to guess at field names and types when chaining tools.
Destructive operation (clear_collection) has no confirmation step, dry-run, or warning. Agents can permanently delete the knowledge base with a single call. Missing error recovery guidance and idempotency hints.
Tool composition is unclear: build_knowledge_base, add_documents, and process_documents all manipulate the knowledge base with overlapping responsibilities. When should an agent call build_knowledge_base vs add_documents vs process_documents? No clear guidance.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 44 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 39 | - | v1 |
获取知识库信息 - Retrieves current knowledge base status including document count and collection metadata
处理文档 - Processes all .txt files in the txt directory, splitting them by headers or fixed size chunks, and returns structured document chunks
搜索 - Searches the knowledge base for documents similar to the given query using vector embeddings and returns results above a similarity threshold
测试查询 - Runs a test query against the knowledge base and displays the response with source documents
No parameter constraints (enums, ranges, regex patterns) defined. Numeric parameters (top_k, max_tokens, threshold) lack min/max bounds. LLMs may pass absurd values (top_k=1000000, threshold=9999.0) that break API calls or cause timeouts.
API keys and credentials (API_KEY in config.py) are hardcoded in source and exposed to agents. Per pattern:secret-injection, credentials must never appear as tool parameters or be embedded in config, use server-side environment variable injection or vault.
No error handling or recovery guidance. Tools have no documented error responses, retryability, or actionable next steps. An LLM receives a failure but cannot reason about how to proceed.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint). Server cannot signal which tools are safe vs risky. Agents have no way to know that clear_collection is destructive without reading the description text.
Descriptions translated from Chinese (e.g., '构建知识库 - Builds...', '生成回答 - Generates...') are inconsistent in English tone and clarity. Some descriptions are under 50 characters, below the rubric baseline of 194 chars average for A+ tools.