Persistent semantic key-value memory MCP server with ChromaDB backend
MemoriousMCP presents well-structured tool definitions with excellent long-form descriptions that include security warnings, usage guidelines, and LLM-specific instructions. All three tools have clear names starting with action verbs (store, recall, forget) and comprehensive parameter documentation. However, there are notable gaps in output schema documentation and error handling guidance. The store() tool returns only {id}, recall() returns {results}, and forget() returns {deleted_ids}, but these output structures lack formal schema documentation. Error handling is minimal, no guidance on what happens if ChromaDB fails, if embeddings are unavailable, or how to recover from partial failures. The descriptions are exceptionally detailed (addressing vector similarity, key canonicalization, security concerns), but this length (1000+ chars per tool) may dilute key decision-making information for LLMs. Parameter schemas are properly typed (string, integer) with descriptions, but lack constraints like min/max length for keys or bounds on top_k.
Delete stored memories that match a query key. IMPORTANT: Deletion operates on short, canonical keys. The LLM MUST issue forget calls using the same concise, embedding-optimized, space-separated key style used to create memories (otherwise relevant memories may not be found). Prefer 1–5 words separated by spaces when requesting deletions. This tool SHOULD be called by the LLM when the user explicitly requests that certain stored information be forgotten or removed (for example: "forget that I live in Paris") or when the assistant decides a memory must be purged because it is incorrect or sensitive. Parameters: - key: concise, canonical, space-separated query text used to find candidate memories to delete. - top_k: number of nearest matches to consider for deletion. Behavior: - Deletion is irreversible; the LLM should confirm with the user when intent is ambiguous before invoking this tool. - The tool returns `deleted_ids` for the memories that were removed.
Retrieve stored memories relevant to a query key. IMPORTANT: To get reliable results the LLM MUST query with the same short, canonical, embedding-optimized keys used at store time. Keys should be compact (1–5 words, space-separated) and represent the core concept — avoid long descriptive queries. If the current user utterance is verbose, the LLM should first map or canonicalize it to an appropriate short key before calling this tool (for example map "I really like listening to jazz music" -> "likes jazz"). This tool SHOULD be called by the LLM when it needs to fetch previously stored facts, personal details, or preferences to inform a response or provide personalized behavior (for example: to recall a user's favorite cuisine before making restaurant suggestions). Parameters: - key: concise, embedding-friendly, space-separated query text used for similarity search. - top_k: maximum number of nearest memories to return. Returns a dict with `results` (memory items including stored value). If nothing matches, `results` is empty.
Output schemas are not formally documented. store() returns {id: string}, recall() returns {results: list}, forget() returns {deleted_ids: list}, but the response structure, field types, and nested object shapes are inferred from code rather than explicitly declared. LLMs cannot reliably extract fields from undocumented outputs.
No error handling guidance. If ChromaDB fails to initialize, embedding functions become unavailable, or vector similarity returns no results, the tools do not document what error codes or recovery steps the LLM should expect. store() and forget() are destructive but lack confirmation/dry-run patterns or explicit warnings about side effects.
Parameter constraints are missing. 'key' has no min/max length, regex pattern, or character restrictions documented in the schema. 'top_k' defaults to 3 but has no min/max bounds in the schema. 'value' has no size limit. Without these constraints, LLMs can pass invalid inputs (e.g., top_k=999999 or empty key='').
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 58 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 59 | - | v1 |
Store a user's fact, piece of information, or preference for later recall. CRITICAL SECURITY WARNING: NEVER EVER store secrets, passwords, API keys, authentication tokens, private keys, or any other sensitive credentials. This storage is NOT secure and should only be used for non-sensitive information like preferences, facts, and general user data. IMPORTANT: `key` MUST be short, canonical, and optimized for embedding/vector similarity lookups. Prefer 1–5 words separated by spaces (for example: "likes jazz", "pref cuisine italian", "lives in paris"). Do NOT use long sentences or paragraphs as keys — put long text into `value` instead. This tool SHOULD be called by the LLM whenever the user states a fact, personal detail, or stable preference that the assistant is expected to remember. Guidelines for the LLM: - Call this tool for user-expressed facts, identity details, or explicit preferences that will be useful later. - Use `key` as a short, consistent, space-separated descriptor across related memories to improve retrieval quality (canonicalize synonyms where possible). - Use `value` for the full text of the fact or preference to be stored and returned on recall; include any extra context inside `value`. Privacy: avoid storing highly sensitive data (passwords, social security numbers, bank details) unless the user explicitly requests secure storage and consents.
Description length and density may harm LLM decision-making. Each tool description exceeds 1000 characters with nested guidelines, security warnings, and usage patterns. While comprehensive and well-intentioned, this verbosity (baseline A+ tools: 194 chars, p90=392) risks token waste and buries the core action in prose. Consider separating security warnings and implementation notes into a separate server-wide documentation.
No idempotency guarantees. store() generates a new UUID for each call, so calling store(key='likes jazz', value='very much') twice creates two separate records with different IDs. recall() is stateless and safe, but forget() may delete different records on retry if embeddings or the corpus change. Agents retrying on transient failures may create duplicates.