MCP Knowledge Server for semantic document search with multi-context support and OCR processing
Knowledge MCP has clear, actionable tool schemas with explicit input definitions and reasonable descriptions. All 11 tools are explicitly registered with proper JSON Schema structures. However, there are significant gaps: (1) no output schemas documented anywhere, (2) parameter descriptions lack depth and formatting constraints, (3) no error handling guidance in tool descriptions, (4) missing tool annotations (readOnlyHint, destructiveHint, idempotentHint despite risk classifications), (5) tool naming is acceptable but not optimally idiomatic (verb_noun pattern mostly followed). The foundation is solid but the execution lacks production polish. Descriptions average ~80 chars, below the 194-char baseline for A+ tools. No evidence of error categorization, recovery guidance, or dependency hints between tools.
Add a document or image to the knowledge base for semantic search
Clear all documents from the knowledge base
Create a new context for organizing documents
Delete a context and all its vectors (documents remain in other contexts)
List all contexts in the knowledge base
Show details of a specific context including its documents
Remove a specific document from the knowledge base
No output schemas documented. Tools define inputs but provide zero guidance on what fields, types, and structure responses contain. LLMs cannot plan downstream tool calls or extract required IDs (e.g., document_id from knowledge-add) without seeing output structure.
Tool annotations missing despite explicit risk classifications. Destructive tools (knowledge-remove, knowledge-clear, knowledge-context-delete) are marked DESTRUCTIVE but carry no destructiveHint=true. Read-only tools lack readOnlyHint=true. Idempotent operations (knowledge-add with same file) are not marked. These annotations are essential for LLM planning and safety guardrails.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-21 | D | 55 | <=2025-11-25 | v2 |
| 2026-03-09 | C | 63 | - | v1 |
Search the knowledge base using natural language query
List all documents in the knowledge base
Get knowledge base statistics and status
Get status of an async processing task
Parameter descriptions lack formatting constraints and are too terse. E.g., 'Path to the document or image file' for file_path does not specify accepted file types, path format (absolute/relative), or size limits. 'Natural language search query' does not hint at complexity (short phrase vs paragraph), language, or length. Descriptions average ~80 chars vs 194-char baseline. Missing 'When to use' and 'Prerequisites' context.
No error handling or recovery guidance in tool descriptions. Destructive operations (knowledge-remove, knowledge-clear, knowledge-context-delete) require a 'confirm' flag but provide no guidance on what the LLM should do if confirmation is omitted or if deletion fails. No categorization of errors as retryable (transient network fault) vs user-fixable (invalid context name) vs fatal (permission denied).
Tool descriptions omit state-mutation clarity. E.g., knowledge-add says 'Add a document' but does not state whether duplicates are allowed, whether async processing updates state immediately or only after completion, or whether the tool is idempotent (re-adding the same file). knowledge-add has an 'async' parameter defaulting to true but provides no guidance on how the LLM should wait for completion or handle task failures.
No documented inter-tool dependencies or chaining guidance. If knowledge-add returns a task_id, the LLM should call knowledge-task-status to wait for completion before calling knowledge-search. However, no tool description hints at this sequence. Similarly, knowledge-context-create must succeed before adding documents to that context, but no dependency is documented.
knowledge-show and knowledge-context-list lack pagination parameters. Descriptions do not mention whether results are capped, what happens if limit exceeds available documents, or if there is a default. No mention of a 'next_cursor' or 'total_count' in responses. For large knowledge bases, returning all documents risks context window exhaustion.
Parameter constraints are incomplete. knowledge-search top_k has min=1, max=50 but no rationale (why capped at 50?). contexts parameter is a comma-separated string, no description of case sensitivity, whitespace handling, or what happens if an invalid context is named. knowledge-context-create name is 'alphanumeric, dash, underscore, 1-64 chars' in description but not enforced in schema (no pattern regex). Enumeration of invalid inputs and format examples missing.
Tool descriptions are generic and lack LLM-optimized language. 'Get knowledge base statistics and status' (knowledge-status) does not explain what statistics are returned, when to call it (e.g., after batch add), or how it differs from knowledge-show. Descriptions do not include 'When to use', 'Prerequisites', or 'Next steps'. This forces LLMs to guess when and why to call each tool.