Robust Long-Term Memory System for AI Companions - Hybrid approach combining ChromaDB (vector) + SQLite (structured) + File backup with semantic search via embeddings, structured metadata queries, automatic backups and data integrity, migration-friendly exports, and scalability to decades of conversations
The server defines 12 tools with explicit schemas and descriptions visible in the source code. Tool naming follows verb_noun conventions (save_memory, search_memories, delete_memory, etc.), which is good. However, there are significant gaps: (1) Most parameter descriptions are minimal (5-15 characters), well below the 72-char baseline and insufficient for LLM reasoning. (2) Output schemas are not documented, the code shows Result, MemoryRecord, SearchResult dataclasses but no explicit schema returned to the MCP client for tools' responses. (3) Error handling is basic; no recovery guidance or categorization (retryable vs fatal). (4) Several tools lack important constraints (e.g., search_memories has no enum for memory_type; importance levels lack documented 1-10 range in parameter descriptions). (5) Parameter consistency issues: some tools use 'importance' (integer), others use 'importance_min/max', naming could be clearer. The dataclasses show internal structure, but there's no evidence of Streamable HTTP transport, the server uses fastmcp with STDIO, which is a hard transport cap. Overall, definitions are above F-level but have enough gaps to score in the C/D range.
Create a full backup of all memories to a JSON file. Returns path to backup file.
Delete a memory record by its ID. Returns success/failure.
Export memories as JSON with optional filtering by memory_type, tags, or importance range.
Retrieve a specific memory record by its ID. Updates last_accessed timestamp.
Get system statistics including total memories, counts by type/tags, importance distribution, and database size.
Import memories from a JSON file. Can merge with existing memories or overwrite.
List all memories with optional filtering by memory_type, importance, or tags. Supports pagination.
Output schemas not documented for MCP client consumption. The code defines internal dataclasses (MemoryRecord, SearchResult, Result) but does not expose what structure is returned to the LLM via the MCP protocol. Without documented output schemas, LLMs cannot plan downstream tool calls or extract the right fields.
Parameter descriptions are too brief (5-15 characters on average). Examples: 'Search query for semantic search' (33 chars), 'The ID of the memory to retrieve' (34 chars). These lack actionable detail about format, constraints, and when to use each parameter. Descriptions should be 50-200 characters and include range/enum constraints inline.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-21 | D | 57 | <=2025-11-25 | v2 |
Restore memories from a previously created backup JSON file. Can merge or overwrite existing memories.
Save or update a memory record in the system. Returns success/failure with reason.
Search memories using semantic search, metadata filtering, or exact matching. Supports filtering by memory_type, tags, importance range, and date range.
Update specific fields of an existing memory record. Only provided fields are updated.
Verify database integrity and repair any issues. Returns status and any repairs made.
Missing enum constraints for multi-valued parameters. 'memory_type' is described as 'Type of memory: conversation, fact, preference, event, task, ephemeral, etc.' but is not declared as an enum in the schema. LLMs may hallucinate invalid types. Same issue with 'match_type' in search results.
No documented recovery guidance in error messages. When operations fail (e.g., memory_id not found, restore fails, integrity check fails), there is no indication of what the agent should try next. Should it search for the ID? Try another backup? The code returns Result with success/reason but offers no recovery paths.
Destructive operations (delete_memory, restore_memories with merge=false, import_memories with merge=false) lack confirmation or dry-run support. An agent could accidentally delete all memories or overwrite with wrong backup without confirmation. No idempotency guarantees stated.
Pagination in list_memories is underspecified. The tool accepts 'limit' and 'offset' but does not document max limit, whether a total_count is returned, or if there's a next_cursor pattern. Large memory stores could return thousands of records, overflowing context. The response schema should clarify pagination.
Parameter naming inconsistencies reduce clarity. 'importance' vs 'importance_min'/'importance_max' uses different patterns across tools. 'tags' sometimes expects an array, sometimes a string. These inconsistencies force the LLM to re-read each parameter description instead of inferring from the name alone.
backup_memories and export_memories both produce file-based outputs but differ: backup returns a path, export requires an output_path parameter. The distinction is unclear, when should an agent use each? The descriptions do not clarify the use case or when one is preferred.