Persistent memory for Claude Code — never lose context between sessions
memory-mcp demonstrates good naming discipline and comprehensive parameter documentation, but suffers from inconsistent schema quality, output documentation gaps, and limited error recovery guidance. All 10 tools have descriptions and follow verb_noun naming conventions (memory_*). Parameters are well-typed via zod with enums where appropriate (MEMORY_TYPE). However, output schemas are not formally documented, tools return text responses without structured field declarations. Error handling is minimal: tools return success/failure text but do not provide recovery guidance (e.g., when a memory is not found). The 'memory_ask' tool calls an external LLM (Haiku) without timeout constraints or fallback behavior. No tool demonstrates idempotency guarantees or confirmation patterns for destructive operations. The server correctly avoids exposing credentials as parameters.
Ask a question and get an answer synthesized from project memories. Like RAG over your project knowledge.
Generate the full consciousness document. This is what gets written to CLAUDE.md.
Manually trigger memory consolidation. Merges duplicates, removes outdated memories, keeps memory sharp.
Delete a specific memory by ID.
Initialize project memory with name and description.
Recall all active memories, optionally filtered by type or tags.
Get all memories related to specific tags/areas. Use to explore a topic in depth.
Output schemas not documented. Tools return generic text responses without structured field definitions. LLMs cannot reliably extract data for chaining or planning downstream operations. Violates pattern:tool requirement that 'tools returning lists should accept page/offset and limit parameters and return a total count or next_cursor'.
External LLM calls (memory_ask, memory_consolidate) lack timeout constraints, retry logic, and fallback behavior. No documentation of expected latency or failure modes. If callHaiku() hangs or returns null, error handling degrades to returning raw memory snippets without proper error classification or recovery guidance.
Destructive operation (memory_delete) lacks confirmation pattern or dry-run option. No idempotency marker or recovery guidance. Description does not state that deletion is permanent or irreversible, which agents need to know before executing.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 59 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 50 | - | v1 |
Save a memory about this project. Records decisions, patterns, architecture, gotchas, progress, or context for future sessions.
Search memories by keyword. Returns ranked results matching the query across content and tags.
Show memory statistics: counts by type, active/archived/superseded, last consolidation.
Error responses are generic text without actionable recovery steps. When memory_ask finds no relevant memories, it returns 'No relevant memories found to answer this question' but does not suggest calling memory_save to create context or memory_related to explore related topics. Violates pattern:recovery-guide.
memory_consolidate and memory_consciousness tools have empty or minimal input schemas ({}) but perform significant operations. No parameters to control scope, filtering, or output format. LLMs cannot customize behavior and may invoke these tools at inappropriate times.
Output field naming inconsistency. memory_recall and memory_search return memory snapshots with fields [id, type, content, tags] but no metadata (created_at, updated_at, superseded_by). Response fields needed for downstream chaining (e.g., passing a memory ID to memory_delete) are present, but temporal context for audit/versioning is missing.
No rate limiting or runaway protection documented. memory_ask calls external Haiku service for every query without throttling. An agent in a loop could generate hundreds of LLM calls per minute, exhausting budget or triggering service limits.
Pagination not implemented. memory_recall, memory_search, and memory_related can return large result sets (no hard cap visible in code). Returning hundreds of memory entries in a single response will exhaust context window and degrade LLM reasoning. Pattern requires pagination with limit enforcement.