Persistent memory for Claude — Ebbinghaus forgetting curve, semantic deduplication, MCP-native
YourMemory has 9 tools with schemas and descriptions present, but significant quality gaps prevent a higher score. All tools have basic descriptions (10-200 chars), and most have input schemas with type definitions. However, descriptions are often terse and lack LLM-guidance on WHEN to use tools vs. alternatives. Parameter descriptions are minimal (e.g., 'User ID' instead of 'Unique identifier for the user, use display name if ID unknown'). Output schemas are not documented, critical for agent chaining. No error handling guidance, tools return errors but do not advise recovery steps. Tool composition has weaknesses: retrieve, observe, and buffer-store overlap in memory management concerns but lack explicit parameter dependency documentation. Schema validation is present but parameter constraints (ranges, enums) are underspecified in descriptions. Naming is verb-first but some names are vague: 'observe' could mean read-only but actually performs fact extraction (side effect). No tool annotations (readOnlyHint, destructiveHint) to signal safety properties to agents. Against baselines: 9 tools avg 14.4 chars per name (low; most under 10 chars with hyphens), avg description 58 chars (below 194 baseline for production tools), avg 4.1 params per tool (matches baseline), but 0% have documented output schemas (vs 100% A+ baseline).
Retrieve the last n verbatim conversation exchanges for injection
Store verbatim conversation exchanges in a buffer for lean-window mode
Trigger memory compaction to merge redundant clusters and apply decay
Delete a single memory by ID
Health check endpoint
List all memories for a user with pagination
Extract atomic factual statements from work output (code, documents, logs) using LLM-based fact extraction
No output schemas documented for any tool. LLMs cannot plan downstream calls or extract chaining IDs (e.g., after calling list-memories, what fields are returned?). Agents lack visibility into response structure.
'observe' is named as a verb but performs fact extraction with side effects (writes to memory). Ambiguous name, LLMs may assume read-only. Should be 'extract_facts' or 'extract_and_store_facts' to signal intent.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint). Agents cannot infer safety properties, e.g., delete-memory should have destructiveHint=true; retrieve should have readOnlyHint=true. This prevents safe automatic retries and composition.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | D | 52 | 2026-07-28+ | v2 |
Retrieve memories using semantic search, BM25, and entity graph
Store a memory fact with optional importance, tags, and metadata
Descriptions lack LLM-guidance on WHEN to use each tool. 'Retrieve memories using semantic search, BM25, and entity graph' (retrieve) does not explain when to call retrieve vs. buffer or list-memories. No dependency hints or alternative guidance.
Parameter descriptions are generic and lack constraints. E.g., 'topK' in retrieve has no range guidance (what if LLM passes 1000?); 'importance' in store has no example (0-1 scale not obvious from '0-1' alone). No enums or ranges in schema descriptions.
No error handling or recovery guidance in tool descriptions. If retrieve returns empty results or delete-memory fails with a foreign key error, LLMs have no guidance on next steps. No examples of error responses or recovery patterns.
Tool composition is weak. 'retrieve' and 'buffer' both provide memory recall but with different semantics (semantic search vs. verbatim). Descriptions do not explain when to use each. 'observe' and 'store' both write memory but with different purposes (fact extraction vs. direct storage). Missing explicit tool-chaining documentation.
'compact' tool lacks clear purpose in descriptions. What does 'merge redundant clusters' mean to an LLM? When should an agent call compact? Only once? Before storing? No guidance.
Parameter names mix camelCase (userId, createdAt) and snake_case (user_id, user_text). Inconsistency confuses agents. Suggest standardizing to snake_case across all parameters.