Semantic memory layer for AI applications. REST API + MCP transport + knowledge graph + autonomous consolidation. Works with 14+ AI clients. Self-host, zero cloud cost.
This MCP server has 12 tools with mostly complete descriptions and schemas. Naming is verb-forward and clear (memory_store, memory_search, memory_delete, etc.). Descriptions are detailed and include explicit 'USE THIS WHEN' guidance that helps LLMs understand when to invoke each tool. Input schemas are present and typed. However, there are notable gaps: output schemas are not documented anywhere in the provided source; error handling patterns are absent (no recovery guidance, no error categorization, no retry logic); parameter descriptions lack specificity on constraints (ranges, formats, enum details in some cases); some parameters have overlapping or confusing purposes (e.g., memory_delete accepts both 'tag' and 'tags', memory_search has both 'n_results' and 'limit' with no clear distinction). The composition logic is sound (tools have single responsibilities and produce results that chain), but execution guarantees (idempotence, atomicity) are not declared. Security considerations like permission checks, audit trails, and rate limiting are not visible in the schema layer. Parameter validation rules (e.g., similarity_threshold 0.0-1.0 for memory_cleanup) are mentioned in descriptions but not enforced via JSON Schema constraints.
Clean up the memory database by removing duplicates and near-duplicate memories. Helps maintain a high-quality knowledge base by consolidating redundant information.
Manage memory consolidation - trigger consolidation, check status, or control autonomous consolidation. Consolidation merges similar memories and removes duplicates to maintain an efficient knowledge base.
Delete a specific memory by its unique content hash identifier - permanent removal of a single memory entry. This is a PERMANENT operation - memory cannot be recovered after deletion. You must have the exact content_hash (obtained from search/retrieve operations). Only deletes the single memory matching the hash.
Explore memory connections and relationships. Find connected memories, shortest paths between concepts, and memory subgraphs for understanding knowledge structure.
Check database health, storage backend status, and retrieve comprehensive memory service statistics. Use this when: User asks 'how many memories are stored', 'is the database working', 'memory service status'. Diagnosing performance issues or connection problems. User wants to know storage backend configuration (SQLite/Cloudflare/Hybrid). Checking if memory service is functioning correctly. Need to verify successful initialization or troubleshoot errors. User asks 'what storage backend are we using'.
No output schemas documented. LLMs cannot plan downstream tool calls or extract return fields without knowing the shape of responses (e.g., what fields memory_search returns, whether results include timestamps or confidence scores). This forces agents to reverse-engineer output shapes from trial and error.
memory_delete has ambiguous, overlapping parameters. It accepts 'content_hash' (delete single item), 'tag' (string), and 'tags' (array) with no clear semantics: does tag delete all memories with that tag? Can you pass both content_hash and tags? Passing multiple delete modes invites undefined behavior. Split into memory_delete_by_hash and memory_delete_by_tags, or document mutually exclusive parameter rules.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 68 | 2026-07-28+ | v2 |
| 2026-03-09 | C | 65 | 1.8.0+ | v1 |
Ingest documents from files or directories into memory. Supports PDF, DOCX, PPTX, and other document formats. Extracts content and stores it with semantic search capabilities.
List and browse stored memories with pagination and optional filtering by tags. Returns all memories categorized with specific tags (OR logic by default) or lists all memories with limit/offset pagination.
Rate memories and retrieve quality metrics. Assess memory relevance and importance for knowledge base curation.
Search stored memories using semantic similarity - finds conceptually related content even if exact words differ. USE THIS WHEN: User asks 'what do you remember about X', 'do we have info on Y', 'recall Z'. Looking for past decisions, preferences, or context from previous sessions. Need to retrieve related information without exact wording (semantic search). General memory lookup where time frame is NOT specified. User references 'last time we discussed', 'you should know', 'I told you before'. THIS IS THE PRIMARY SEARCH TOOL - use it for most memory lookups.
Get comprehensive memory service statistics including cache performance, consolidation metrics, and storage utilization. Provides insights into service health and performance characteristics.
Store new information in persistent memory with semantic search capabilities and optional categorization. USE THIS WHEN: User provides information to remember for future sessions (decisions, preferences, facts, code snippets). Capturing important context from current conversation. User explicitly says 'remember', 'save', 'store', 'keep this', 'note that'. Documenting technical decisions, API patterns, project architecture, user preferences. Creating knowledge base entries, documentation snippets, troubleshooting notes. THIS IS THE PRIMARY STORAGE TOOL - use it whenever information should persist beyond the current session.
Update metadata or tags of an existing memory by its content hash without modifying the core content. Supports updating tags, type classifications, and other metadata attributes.
memory_search has redundant parameters 'n_results' and 'limit' with no documentation of their difference or precedence. LLMs will pass both and expect undefined behavior, or pass neither and get unexpected result counts. Pick one name ('limit' is more standard) and document its behavior.
Error handling is completely absent. No tool description includes recovery guidance ('if memory_search returns no results, try broader query or use memory_list'), no error categorization (retryable vs fatal), no examples of error messages the LLM should expect. Agents will not know how to handle failures.
Destructive tools (memory_delete, memory_cleanup) lack confirmation/dry-run patterns. memory_cleanup supports a dry_run flag (good), but memory_delete does not. Agents cannot preview what will be deleted before executing. Implement a dry-run or confirmation step for all destructive operations.
Parameter validation rules are mentioned in prose descriptions but not enforced in JSON Schema. E.g., memory_cleanup documents 'similarity_threshold (0.0-1.0)' in description, but schema has no minumum/maximum constraints. LLMs will not read prose constraints, use JSON Schema 'minimum', 'maximum', 'pattern', 'enum' fields.
memory_stats and memory_health have empty input schemas ({}). While these are read-only operations with no inputs, the descriptions do not explain what fields the responses include. Agents cannot know whether to expect cache_hits, memory_count, storage_size, or other metrics.
memory_list 'tags' parameter accepts both array and string types, but behavior is undefined. Does 'tags': 'python,ai' work? Does 'tags': ['python', 'ai'] work? Must both work? Specify the canonical format and reject the other with clear error messaging.
Idempotence and retry semantics are not declared. Can memory_store be called twice with identical input and expect the same memory to be returned (or an error 'already exists')? Or does it create a duplicate? Agents need to know what is safe to retry.
memory_ingest 'metadata' parameter includes 'tags' that accept both array and string, mirroring the confusion in memory_store. Standardize on one format (recommend array) and document it consistently across all tools.