A memory system for AI conversations with semantic search, PostgreSQL storage, and optional encryption
Memory MCP-CE has solid naming conventions (all tools start with action verbs: store, retrieve, get, update, delete, list), and all 7 tools include descriptions. However, there are significant gaps in schema completeness, parameter validation clarity, and output documentation. Input schemas are present and well-formed with proper type declarations and parameter descriptions. The descriptions are moderately detailed (100-180 chars typical), exceeding the 10-char minimum but often lacking strategic context about WHEN to use the tool vs. alternatives. Output schemas are NOT documented in the tool definitions, the server returns complex objects (embeddings, related memories, state JSON) but LLMs cannot infer the response structure without explicit schema. Error handling is absent from tool definitions; there is no guidance on what errors might occur or how to recover. The tools follow composition best practices (single responsibility, clear chains via ID returns), but miss the formal patterns around error classification, recovery guidance, and idempotency markers.
Delete a memory by ID. Removes backlinks from related memories and returns confirmation.
Get details of a specific memory by ID. Returns full memory object with labels, source, and related memories.
Get trending labels in the memory namespace based on recent usage and decay. Returns top labels with usage counts.
List all memories in the namespace with optional filtering by labels or source. Returns paginated list.
Retrieve memories by semantic similarity to a query. Returns list of matching memories with similarity scores and optional related memories.
Store a memory with optional labels, source, and related memories. Returns memory ID, embedding dimension, and optional timezone/performance info.
Output schemas are not documented anywhere in tool definitions. Tools return complex nested objects (memory state with embeddings, related memories, timestamps) but LLMs have no formal schema to understand response structure. This forces LLMs to guess what fields are present and makes chaining unreliable.
No error handling or recovery guidance in any tool description. Tools can fail (e.g., delete_memory when memory not found, store_memory on encryption errors, retrieve_memories on embedding service unavailability) but descriptions do not mention error cases or what to do next. This leaves LLMs unable to self-correct.
Inferred effective spec: 2026-07-28+.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | B | 74 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Update the content and/or labels of an existing memory. Returns updated memory details.
Destructive operation delete_memory has no confirmation or dry-run pattern. The code shows the tool directly deletes and cascades cleanup to backlinks, but there is no safety mechanism (pre-flight check, confirmation step) to prevent accidental deletions by an LLM.
Parameter 'limit' appears in retrieve_memories (default 10), list_memories (default 50), and get_trending_labels (default 20) but there is no documented maximum or enforced cap. Unbounded or very high limit values could cause context window exhaustion. No guidance on what happens if limit exceeds available memory.
Idempotency not declared. store_memory, update_memory, and delete_memory are state-modifying operations, but tool descriptions do not state whether they are idempotent. An LLM retrying a failed update or store could produce duplicates or unexpected state changes.
Parameters 'days' in retrieve_memories and 'offset' in list_memories lack validation documentation. No minimum/maximum stated. LLM could pass days=-1 or offset=-100 without guidance on valid ranges.
Tool descriptions do not explain WHEN to use one tool vs. another. E.g., retrieve_memories vs. list_memories vs. get_memory all read from the store, when should an LLM choose which? Descriptions lack strategic context.
store_memory description mentions 'optional timezone/performance info' in return but does not specify output schema (which fields, types, when they appear). This is vague and makes LLM response parsing unreliable.