Temporal knowledge graph MCP server for storing and querying architectural decisions, patterns, and failures with hybrid search and gap detection
Mixed quality across 10 tools. Naming is generally strong (verb-based, mostly clear). Descriptions are present but often too brief (20-50 chars) to be LLM-optimal (target 50-200 chars). Input schemas are well-defined with proper type constraints and descriptions for parameters. However, critical issues: (1) add_decision_json bypasses normal parameter validation by accepting a raw JSON string, this is documented as a workaround for a tool-call parsing issue that should be fixed at the framework level, not worked around; (2) output schemas are NOT documented, tools return responses but LLMs have no formal specification of response structure; (3) error handling is minimal, most tools lack recovery guidance or actionable error messages; (4) two tools (find_influential_patterns, find_knowledge_communities) appear to be registered only in server_fastmcp.py, not visible in the provided source, capping their scores at 50; (5) no tool has toolAnnotations (readOnlyHint, destructiveHint, idempotentHint) despite several being clearly destructive (add_decision, add_pattern, add_failure) or read-only (query_decisions, find_related, detect_gaps, get_timeline). The clamping-and-capture mechanism for over-length inputs is thoughtful, but truncation with a marker and async overflow logging adds complexity without being formally declared in parameter constraints.
Record an architectural decision with full context and reasoning.
Record an architectural decision from a single JSON string payload. Use this variant when a normal add_decision call fails with 'rationale Missing required argument' — that is a tool-call parsing issue with long multiline parameters; passing one JSON string avoids it.
Document what didn't work and lessons learned.
Store successful implementation pattern.
Run NetworkX structural gap analysis. By default excludes SEMANTICALLY_SIMILAR edges so bridges and isolated nodes reflect real structural gaps rather than embedding noise.
Find most connected/influential patterns using PageRank algorithm. Returns patterns ranked by their influence in the knowledge graph, based on how many other nodes reference them.
No output schemas documented. Tools return responses but LLMs have no formal specification of return types, fields, or structure. This forces LLMs to infer output shape, leading to hallucinated field access and failed downstream tool chaining.
add_decision_json is a workaround tool that accepts a raw JSON string to bypass parameter parsing issues in the framework. This is a symptom of a deeper problem: the framework's parameter serialization is broken for long multiline text. The tool should not exist, instead, fix the root cause in fastmcp or fall back to a different parameter encoding strategy.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 57 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 45 | - | v1 |
Detect communities of related knowledge using connected components. Groups patterns that are strongly connected to each other, revealing clusters of related knowledge.
Find related knowledge nodes via graph traversal. By default, only structural edges (RELATES_TO, SOLVES, IMPLEMENTS, etc.) are traversed. The noisy SEMANTICALLY_SIMILAR edges from embeddings are excluded unless include_similar=True.
Get temporal view of how knowledge evolved over time.
Search decisions using hybrid graph+vector search. limit: max results (1-50, default 15).
Missing tool annotations (readOnlyHint, destructiveHint, idempotentHint). add_decision, add_pattern, add_failure are destructive writes; query_decisions, find_related, detect_gaps, get_timeline are read-only. Agents cannot infer safety profile without annotations.
Tool descriptions are too brief (20 - 75 chars) to be LLM-optimal. Target is 50 - 200 chars. E.g., 'Record an architectural decision with full context and reasoning' (62 chars) lacks guidance on WHEN to use this vs alternatives, and doesn't explain what data structure is returned.
Minimal error handling and no recovery guidance. Most tools return simple dicts without error codes or actionable messages. E.g., if query_decisions finds nothing, it returns silence instead of 'No decisions matched your query. Try broadening the search or filtering by a different date range.'
Two tools (find_influential_patterns, find_knowledge_communities) are only visible in server_fastmcp.py; their actual implementation is not provided in source. Cannot verify input/output schemas, error handling, or internal logic. Capping these at 50.
Clamping-and-capture mechanism (truncating long inputs and logging full text to overflow_capture.jsonl) is not formally declared in parameter constraints. LLMs do not know that 'rationale' input will be silently truncated if it exceeds the max_length. Add explicit truncation notice to parameter descriptions.