Persistent project memory for Claude Code — track architectural decisions and query them conversationally
The server has 7 tools with explicit registration, complete input schemas using Zod, and clear descriptions. However, there are significant gaps: most parameter descriptions are minimal (averaging ~30-50 chars vs. the 72-char baseline), many optional parameters lack explanation of when to use them, output schemas are NOT documented (only inferred from code), and error handling is generic without recovery guidance. Tool names follow verb_noun convention (save_decision, query_memory, list_recent, update_decision, get_stats, sync_claude_md, review_session) and are reasonably descriptive. The server demonstrates awareness of composition patterns (separate read vs. write tools) and addresses a genuine problem domain (persistent decision memory). But compared to production baselines, the implementation lacks depth in parameter guidance, output documentation, and error classification.
Get project memory statistics (call at session start for context)
List recent decisions from current or past sessions
Search for decisions using natural language
Generate a contextual summary of this session's architectural decisions for Claude
Save an architectural decision made during this session
Sync active decisions into CLAUDE.md (auto-generates a decisions section)
Update or deprecate an existing decision
Output schemas are not documented. Tool descriptions state what each tool returns in prose (e.g., 'Decision saved (ID: X)' for save_decision), but there is no formal schema definition visible to the LLM. This forces LLMs to infer the structure of response fields and makes it hard to chain tools or extract specific fields reliably.
Parameter descriptions are sparse and lack actionable guidance. For example, save_decision has 'alternatives' (array of strings) with description 'What else was considered', no guidance on format, count, or when this is critical vs. optional. Similarly, 'context' is described as 'What prompted this decision' but doesn't explain when it's essential for decision tracking. compare to baseline of 72 chars avg, these average ~30-40 chars.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 53 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 53 | - | v1 |
Error handling is generic and non-actionable. In src/server.ts, all tools catch errors with 'Error: {error message}', no classification (retryable vs. fatal), no recovery guidance, no suggestions for follow-up actions. Example: if query_memory finds no results, it returns 'No decisions found matching: "<query>"' with no hint to try broader search, alternative queries, or check session filtering.
get_stats and sync_claude_md have minimal descriptions. get_stats: 'Get project memory statistics (call at session start for context)', tells WHEN but not WHAT is returned. sync_claude_md: 'Sync active decisions into CLAUDE.md (auto-generates a decisions section)', unclear what 'active' means, whether this overwrites existing content, or what fields appear in CLAUDE.md.
Mutation tools (save_decision, update_decision, sync_claude_md) lack idempotency guidance and dry-run support. If an LLM retries save_decision, does it create a duplicate? If update_decision is called twice with the same args, is it safe? No confirmation or dry-run pattern documented.
No pagination or result limits documented for query_memory and list_recent. query_memory defaults to limit=5, list_recent to limit=10, but no max cap, no documentation of what happens if the query matches 1000 decisions, and no next_cursor or offset support visible.