A memory management server using FastMCP that records, retrieves, and manages user-specific information with semantic search capabilities powered by Qdrant vector database and Google Gemini embeddings
Memory Plus has 7 tools with mixed quality. Tool definitions are present and registered via FastMCP, but critical gaps exist: naming conventions are inconsistent (e.g., 'set_whether_to_annonimize' has a typo and is awkwardly named), several tools lack proper input validation descriptions, output schemas are not formally documented, and error handling guidance is minimal. The 'record' tool has excellent precondition and decision documentation, and 'retrieve'/'recent' have good descriptions explaining when to use them. However, 'set_whether_to_annonimize' is poorly named (should be 'set_anonymization' or 'configure_anonymization'), and 'visualize_memories' lacks guidance on when/why an agent should call it. No tools declare their output schema structure (Pydantic models or JSON schemas), making it difficult for LLMs to plan downstream operations. Parameter descriptions are mostly present but lack constraint documentation (e.g., 'metadata' dict fields are suggested but not validated). Overall readability is reasonable, but the server falls short of production-grade tool composition.
Delete a memory entry by its ID.
Retrieve the most recently recorded memory entries, up to `limit` items. The assistant should automatically call this function when composing responses that benefit from the freshest context— such as referencing what the user just said, recent preferences, or newly provided details. Each returned entry contains: - `content`: The stored memory text - `metadata`: Associated metadata (e.g., timestamp, category, context source) Returns: A list of memory dictionaries ordered from newest to oldest.
This tool is invoked automatically when the assistant detects enduring, user-specific information—such as stable preferences, personal background facts, or recurring discussion topics—that can enrich subsequent responses. - Trigger: Automatically invoked when the assistant detects a stable user detail (e.g. preference, background fact, recurring topic). - Precondition: - The assistant SHOULD first call `retrieve(content, top_k)` to fetch top_k similar memories. - If resource://recorded_memory_categories has not been called in the current conversation, the assistant should call it first to understand which memory categories already exist. The assistant should prefer using existing categories from the resource when constructing the metadata for the new memory. - Decision: - If any retrieved memory has similarity really high, use `update(memory_id, new_content, metadata)`; - Otherwise, use `record(content, metadata)`. Returns a list of newly assigned memory IDs.
Typo in tool name: 'set_whether_to_annonimize' contains 'annonimize', should be 'anonymize'. This is a critical naming defect that will confuse LLMs and is unprofessional.
No explicit output schemas documented for any tool. Tools return List[int], List[Dict], or unspecified objects, but LLMs cannot infer field names or types. For example, retrieve() returns 'list of memory entries' with 'content' and 'metadata' fields, but the exact structure (field types, required vs optional) is not formal.
delete() tool lacks error handling guidance and recovery instructions. Description is too brief (36 chars) and does not warn that deletion is irreversible or explain what to do if the memory_id is not found.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 53 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Retrieve up to `top_k` memory entries that are most semantically similar to the given `query`. This tool enables contextual awareness by allowing the assistant to recall relevant user-specific information previously stored in memory—such as preferences, background details, ongoing projects, or recurring topics—based on semantic similarity to the input query. The assistant should automatically call this tool when additional context is needed to improve understanding or response quality, even without explicit user instruction. To avoid unnecessary or irrelevant memory retrievals, the assistant should first check `resource://recorded_memory_categories` to understand what kinds of memory categories are currently available. Returns a list of memory entries, each containing: - `content`: The distilled memory text (as originally stored). - `metadata`: Structured metadata describing the memory, including fields such as `category`, `tags`, `timestamp`, `source`, and other contextual details. Returns: List[Dict[str, Any]]: A list of memory entries ordered from most to least relevant.
Set whether to anonimize the content.
Update an existing memory with new content and metadata. - Trigger: When there is an old memory that is similar to the new content, which can be checked whenever there is a new memory recorded. - Precondition: - The assistant should first retrieve the existing memory with `retrieve(content, top_k)` to verify it matches the memory_id intended for update.
Generate a visualization of stored memories using clustering and dimensionality reduction.
Parameter 'metadata' in record() and update() tools is documented as a suggested dict structure in prose, but not validated as a Pydantic model. LLMs may pass invalid metadata (missing required fields like 'category', or extra fields). Validation rules and enforcement are missing.
visualize_memories() tool lacks context for agent decision-making. Description does not explain WHEN an agent should call it, what the output format is (file path, URL, raw data?), or how it fits into the memory retrieval workflow. Is this for user reporting, debugging, or agent introspection?
No tool declares its output schema in the docstring or as a formal type hint. LLMs cannot plan multi-step operations (e.g., knowing that retrieve() returns memory_id so that update() can use it) without explicit schema documentation.
set_whether_to_annonimize() lacks mutation warning and side-effect clarity. Description does not warn that calling this modifies state (writes to a file), does not explain what happens if the file write fails, and does not document when this configuration takes effect.