Persistent local memory MCP server for Claude Code, Codex CLI, Cursor and any MCP client: 74 tools — temporal knowledge graph, procedural memory, episodic memory, AST codebase ingest, pre-edit guard, auto-consolidating error capture
The server defines 20 tools with basic schemas and descriptions, but quality is inconsistent. Tool names follow verb_noun convention (memory_recall, memory_save, memory_update, etc.), which is good. However, descriptions are often vague or missing critical context about when to use each tool vs. alternatives. Input schemas are visible but lack depth: most parameters have type and description, but lack constraints (enums, ranges, patterns). Output schemas are not documented anywhere in the provided source. Parameters accept free-form strings where enums would be safer (e.g., memory_type accepts 'decision, solution, observation' as plain text without an enum constraint). Error handling guidance is absent, tools don't explain what to do if a query returns no results, or how to recover from failures. The server lacks sophistication in parameter validation, output shaping, and LLM-guidance patterns.
Consolidate related memories into higher-level insights and patterns
Permanently delete a memory record
Export memories in JSON, CSV, or Markdown format
Extract and save key learnings from a coding session
Archive or delete memories matching criteria to manage retention and privacy
View version history of a specific memory record
Enqueue a tool observation for deferred extraction
No output schemas documented. Tools return data structures that are not formally described in the provided source. LLMs cannot plan downstream tool calls without knowing what fields to expect.
Parameter constraints not enforced via enums. memory_save accepts 'type' as a free-form string ('decision, solution, observation') with only text description. Should define as enum: ['decision', 'solution', 'observation']. Similar issue in memory_export format and self_rules action parameters.
Destructive operations (memory_delete, memory_forget) lack confirmation or dry-run guidance. No description mentions whether the operation is reversible or if a confirmation step exists. Agents could accidentally delete memories without safeguards.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-21 | D | 57 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 23 | - | v1 |
Retrieve memories from temporal knowledge graph, procedural memory, and episodic memory with BM25 scoring and 3-level progressive disclosure
Create or update relationships in the knowledge graph between memories
Save persistent memories with type (decision, solution, observation), project, tags, importance level
Search memories by tag with fuzzy matching
Get statistics on memory usage: total records, storage size, types distribution
View temporal history of memories organized by date/time with decay scoring
Update an existing memory record with new content or metadata
Log error patterns for self-improvement pipeline
Generate insights from error patterns and consolidate into rules
Analyze recurring patterns in errors and coding behavior
Perform reflection on performance and decision-making patterns
View or update self-improvement rules (SOUL) learned from errors
Get contextual rules and guidelines for current task
Error handling descriptions absent. Tools do not explain what happens on failure (e.g., 'query returns no results', 'memory_id not found', 'invalid session_id'). No guidance on retry strategies or recovery steps.
Weak description clarity on tool disambiguation. memory_recall, memory_search_by_tag, and memory_timeline all retrieve memories but descriptions do not clearly explain WHEN to use each vs. the others. LLMs may pick the wrong tool.
Parameter ranges and format constraints missing. 'days' and 'days_old' parameters lack min/max bounds. 'query' and 'pattern' parameters lack length limits or regex guidance. LLMs may pass absurd values (days=999999) or malformed strings.
No pagination guidance. memory_recall and memory_search_by_tag do not document how many results they return or if pagination is supported. Large result sets could blow context windows.
Tool composition clarity weak. No descriptions explain chains (e.g., 'extract_session' + 'observe' + 'consolidate'). LLMs must infer the intended workflow.
Vague descriptions for self_* tools. Descriptions like 'Generate insights from error patterns and consolidate into rules' and 'Perform reflection on performance and decision-making patterns' are too abstract. What does 'consolidate' do? What fields does 'reflect' return?