Trusted memory system for long-lived coding agents backed by SQLite + FTS5, with built-in truth governance for managing contradictions, staleness, and memory lifecycle.
The agentmem server demonstrates solid tool design with clear naming, comprehensive descriptions, and well-structured schemas. All 13 tools follow verb_noun naming conventions and are accompanied by descriptive text explaining WHAT they do and WHEN to use them. Input schemas are present and typed for all tools. However, output schemas are entirely undocumented, the code shows no documented return types, field structures, or examples. This is a significant gap for agent composition, as LLMs cannot plan downstream tool calls without knowing what fields to expect. Additionally, error handling guidance is minimal; tools provide no recovery suggestions or error classification. Parameter descriptions are generally good but lack explicit constraints (ranges, regex patterns) that would help LLMs avoid invalid inputs. The server also lacks confirmation patterns for destructive operations (delete_memory, while warned against, still accepts a bare ID with no confirmation step).
Store a new memory. Use when something is worth remembering: a preference, fix, decision, or procedure.
Permanently delete a memory by ID. Prefer deprecate_memory for memories that were once true but are no longer. Only delete memories that were created in error or contain incorrect information.
Mark a memory as deprecated. It will be excluded from search/recall but kept for history. Use when a rule or fact is no longer true.
List all memories, optionally filtered by type. Returns memories sorted by most recently created. Use to browse what the memory system knows about a topic or to audit stored rules.
Load the most recent session state. Call this at the start of a conversation to pick up where the last instance left off.
Detect contradictions between active memories. Returns pairs of memories that assert and negate the same topic.
Output schemas completely undocumented. Code shows Tool definitions for inputs but no documented return types, field structures, or examples. LLMs cannot determine what fields search_memory, recall_memory, list_memories, and memory_health return, making downstream tool composition difficult.
No error handling guidance. Tools provide no recovery suggestions (e.g., 'Memory not found. Try search_memory() to find the ID.' or 'Confidence must be 0.0-1.0'). Raw errors force LLMs to guess at next steps.
Destructive operations (delete_memory) lack confirmation or dry-run support. While the description warns 'Prefer deprecate_memory', there is no programmatic safeguard preventing accidental deletion. A confirmation_required pattern would prevent catastrophic mistakes.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | B | 72 | 2026-07-28+ | v2 |
Run a health check on the memory system. Returns: score (0-100), conflict count, stale count, status distribution. Use to audit memory quality.
Promote a memory's trust level: hypothesis -> active -> validated. Use when evidence confirms a memory is true.
Get the most relevant memories for a topic, fitted to a token budget. Use at the start of a task to load context.
Save current session state before conversation ends or context compresses. Capture: what's in progress, what's blocked, what's done, decisions made. The next agent instance loads this automatically.
Full-text search across all active memories. Returns results ranked by relevance, trust status, and recency. Deprecated and superseded memories are excluded automatically.
Replace an old memory with a new one. Old memory is marked superseded and linked to the replacement.
Update the title, content, tags, or confidence of an existing memory. Use when a rule changes, a fix gets refined, or new context applies to an existing memory.
Parameter constraints are implicit, not explicit. 'confidence' is described as '0.0 to 1.0' in text but has no minimum/maximum in the schema. 'stale_days' in memory_health has no bounds. LLMs may pass 2.5 or -100 without validation guidance.
recall_memory and list_memories support pagination (limit parameter) but lack offset/cursor and total_count in documented outputs. Without knowing total_count, agents cannot estimate whether more results exist.
recall_memory description is terse (84 chars) and lacks guidance on what 'token budget' means or how max_tokens is enforced. Is it estimated tokens or strict? Does it truncate results or fail if over budget?