Exposes AgentMem OS memory capabilities as Model Context Protocol (MCP) tools so any MCP-compatible client (Claude Desktop, Cursor, etc.) can use persistent 4-tier memory with a single config entry.
AgentMem OS exposes 6 tools with complete JSON schemas and detailed descriptions (avg 180 chars). Naming follows verb_noun convention (save_, recall_, consolidate_, get_, list_). However, critical gaps exist: (1) descriptions lack actionable error guidance and recovery paths; (2) no tool annotations (readOnlyHint/destructiveHint) despite clear read/write semantics; (3) output schemas are not documented, LLMs cannot predict response structure for downstream chaining; (4) parameter descriptions are verbose but lack format constraints (enums, ranges, patterns); (5) no confirmation/dry-run pattern for destructive operations (save_memory, consolidate_session). Tool composition is sound (each does one thing), but error handling is absent from visible code.
Distill a session's turns into dated, atomic, cited semantic facts (Consolidation v2): schema-constrained local-LLM extraction, entity linking through the knowledge graph (cross-lingual ALIAS_OF), and per-fact supersession judgment — new facts can mark old ones as history, never deleting them. These facts are what recall_memory surfaces first. Requires a local Ollama server; fails loudly (with zero writes) when the LLM is unreachable.
Return the entity relationship subgraph for a session. Shows entities extracted via spaCy NER, co-occurrence edges with weights, and BFS-reachable neighbors up to 2 hops from a seed entity. Useful for understanding what concepts the agent knows about.
Return mined procedural memory patterns for a session. Patterns are (trigger→action) pairs extracted from conversation history with confidence scores. Examples: bug_report→explain_solution (0.82), feature_request→provide_implementation (0.75). Use these to predict what kind of response is most useful.
List all stored sessions with their metadata: turn count, token usage, creation time, and parent branch info. Useful for session management and understanding what conversations are stored in memory.
Retrieve relevant memory context for a query from all 4 memory tiers: episodic (exact turns), semantic — dated atomic FACTS first (distilled by consolidate_session, supersession-aware, with change history), then raw-turn semantic search as provenance/fallback — entity KG (relationship graph), and procedural (behavioral patterns). Returns a structured XML context block ready to prepend to any LLM prompt.
No output schemas documented. LLMs cannot predict response structure (fields, types, nesting) for downstream tool chaining. E.g., recall_memory returns 'structured XML context block' but no schema is visible, agents cannot extract session_id, entity_ids, or fact citations for follow-up calls.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint). save_memory and consolidate_session are clearly destructive (WRITE risk), but MCP clients cannot detect this from the tool definition. Agents cannot reason about retry safety or side effects.
Error handling absent from visible code. No recovery guidance, error classification, or actionable messages. E.g., consolidate_session 'fails loudly when Ollama is unreachable', but no error response structure is defined. LLMs cannot distinguish retryable vs fatal failures.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | D | 59 | 2026-07-28+ | v2 |
Persist a conversation turn into AgentMem OS's 4-tier memory hierarchy. Triggers importance scoring, entity KG ingestion, and background consolidation. Use this after every user/assistant exchange to build persistent memory.
Parameter descriptions lack format constraints. E.g., model_window defaults to 128000 but no min/max bounds stated. min_confidence (0.0 - 1.0) is documented in description text, not as JSON Schema constraints. LLMs cannot validate inputs before calling.
No confirmation/dry-run pattern for destructive operations. save_memory and consolidate_session modify persistent state but offer no preview, undo, or confirmation step. Agents cannot safely explore consequences before committing.