Persistent memory system for Claude Code — hybrid BM25 + vector search, LLM-driven structuring, automatic clustering
This server implements a coherent persistent memory system with 4 well-named tools that follow verb-noun patterns (memory_search, memory_save, memory_validate, memory_stats). All tools have descriptions (average ~80 chars, baseline 194 chars) and input schemas with type definitions. However, descriptions are below the 50-200 char LLM-optimized range, parameter constraints are present but sparse (e.g., confidence has min/max bounds, type/domain use enums), and output schemas are not documented in source. Error handling is not visible in the tool definitions themselves, no recovery guidance or error classification patterns. The tool composition is sound (each does one thing) and naming is consistent and clear. Missing: output schema documentation, comprehensive parameter descriptions, and visible error handling patterns.
Save a new persistent memory. Use to record important patterns, decisions, bug fixes, user preferences, etc.
Search persistent memories (hybrid BM25 + vector semantic retrieval). Use when you need to recall previous context, patterns, decisions, or bug fix records.
View memory system statistics: total memories, type distribution, domain distribution, cluster status, etc.
Validate whether a memory was helpful. Helpful increases confidence by +0.1, unhelpful decreases by -0.05.
Output schemas not documented. Tool descriptions mention what is returned (e.g., 'Search results', 'statistics') but do not specify the structure of the response object, fields, data types, or pagination. LLMs need this to plan downstream tool calls and extract data correctly.
Parameter descriptions are terse (e.g., 'Search query', 'Memory content to save', 'Memory ID'). They lack context on WHEN to use these parameters, WHAT happens with different values, or dependencies between parameters. Baseline for param descriptions is ~72 chars; these average ~40 chars.
Tool descriptions lack explicit guidance on WHEN to use each tool. For example, memory_search says 'Use when you need to recall previous context...' which is good, but memory_validate and memory_stats lack such motivational context. Descriptions should answer: What does it do? When should the LLM call it? What does it return?
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-21 | F | 11 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 57 | - | v1 |
No visible error handling patterns in tool definitions. Error responses should tell the LLM what to do next (e.g., 'Memory not found. Try a different search query.'). No evidence of error categorization (retryable, user-fixable, fatal) or recovery guidance in source.
memory_validate parameter 'is_valid' lacks clarity. The description says 'Whether the memory was helpful (true=helpful, false=not helpful)' but does not explain the side effects: +0.1 confidence for helpful, -0.05 for unhelpful. This is a mutating operation whose consequences are underdocumented.
memory_search limit parameter defaults to 5 but lacks min/max constraints in schema. Baseline guidance: numeric parameters should specify range (e.g., limit 1 - 100). Unbounded or loosely bounded numbers allow LLMs to pass absurd values that may break memory retrieval or waste tokens.
memory_save confidence parameter describes range as '0.3 - 0.9' but schema only shows min/max without explaining why these bounds exist or what confidence means semantically. Descriptions should be actionable: 'Confidence (0.3 - 0.9): how certain you are this memory is accurate. Higher values boost recall priority.'