MCP server providing knowledge base search, retrieval, evidence extraction, and graph traversal tools for WEKA documentation with Neo4j knowledge graph integration
WEKA Docs Matrix MCP server demonstrates solid definition quality with comprehensive tool descriptions and well-structured parameter schemas. All 11 tools have explicit descriptions (ranging 100 - 500+ chars) that explain WHAT the tool does and WHEN to use it. Parameter schemas are consistently defined with type information and descriptions. However, several tools lack explicit output schemas in the source code, and some parameters use generic 'object' types for nested structures without detailed documentation. Tool naming follows verb_noun convention clearly (search_, kb.read_, kb.expand_, graph.expand, etc.), and the descriptions include sophisticated guidance ('IMPORTANT: You MUST use the exact passage_id UUID returned by kb_search') that guides LLM behavior well. Error handling is present but not comprehensive, no explicit error recovery guidance in most tool descriptions. The server manages per-tool token budgets and includes careful parameter constraints (max_tokens=800, max_snippet_chars=500, etc.), showing design sophistication. Output pagination and cursor support are documented in kb.search description. Overall, this is a B+/A- effort: definitions are professional and LLM-aware, but output schemas could be more explicit in code, and error paths lack recovery guidance.
Get child sections in the documentation hierarchy. Returns immediate children with edge metadata.
Get structural summary of a section node. Returns type, properties, and metadata.
Expand graph neighborhood around a seed section. Returns related sections with edges. Use for follow-up navigation when user asks about related content near a specific section.
Get parent sections in the documentation hierarchy. Returns immediate parents with edge metadata.
Find shortest paths between two sections in the knowledge graph. Explains how concepts/sections are related.
Use to expand around the last excerpt for a passage_id from kb_search. IMPORTANT: You MUST use the exact passage_id UUID returned by kb_search. Do not use for discovery; use kb.search first. Returns a bounded expansion before/after the last excerpt window.
Output schemas are imported/referenced but not fully documented in source code. Tools declare KB_SEARCH_OUTPUT_SCHEMA, GRAPH_EXPAND_OUTPUT_SCHEMA, etc., but LLMs see only partial schema structure. This prevents validation of response completeness and forces LLMs to infer output structure.
Generic 'object' type used for 'scope', 'filters', 'options' parameters without nested schema documentation. LLMs cannot infer which fields are accepted or required within these objects. Example: kb.search accepts 'filters' as object with undocumented fields (doc_tag, path_prefix, updated_after mentioned in description but not in schema).
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 59 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 17 | - | v1 |
Use to extract minimal quotes that answer a question from known passage_ids. IMPORTANT: You MUST use the exact passage_id UUIDs returned by kb_search. Do not use for discovery; use kb.search to get passage_ids first. Returns at most max_quotes short quotes with citations.
Use when you already have a passage_id from kb_search and need a bounded excerpt. IMPORTANT: You MUST use the exact passage_id UUID returned by kb_search (e.g., '022160622467452ba0ccf501cbf771c1'), NOT a constructed document path. Do not use for discovery; use kb.search first. Returns a capped excerpt (default 300 tokens, max 800, 32KB per call). If you need more context, call kb.expand_excerpt.
Use when you want the best evidence in one call. This runs kb.search then kb.extract_evidence, returning quotes only. Do not use when you need exploratory graph traversal.
Use when you need a short list of candidate passages. Do not use when you already have passage_ids; use kb.read_excerpt or kb.extract_evidence instead. Returns at most page_size results with previews and scratch URIs (no full text). If you need more text, call kb.read_excerpt or kb.expand_excerpt. Defaults: top_k=5, page_size=5, max_snippet_chars=280, max_per_doc=1. Max: top_k=20, page_size=20, max_snippet_chars=500.
Search documentation using hybrid retrieval (vector + graph). Returns answer_markdown and answer_json with evidence, confidence, and diagnostics. Supports backward compatibility with legacy verbosity tokens.
Two search tools exist (kb.search and kb.retrieve_evidence, plus search_documentation) with overlapping functionality. Descriptions lack clear guidance on WHEN to use each. LLMs must reason through subtle differences ('retrieve_evidence is convenience wrapper' vs 'search_documentation uses hybrid retrieval'). Consider consolidating or adding explicit decision tree.
Minimal descriptions for graph traversal tools (graph.describe: 48 chars, graph.parents: 57 chars, graph.children: 54 chars). These fall below the recommended 10 - 200 char range for LLM interpretation. No guidance on use cases, typical graph depths, or computational costs.
Error handling is not documented in tool descriptions. No guidance on what errors LLMs should expect, whether they are retryable, or how to recover. Pattern: recovery-guide would strengthen resilience.
Parameter 'retrieval_depth' (kb.retrieve_evidence) lacks explanation. Defaults to 60, max 150, but no guidance on what this parameter controls or computational trade-offs. LLMs cannot reason about whether to adjust it.