Self-hosted hybrid memory & retrieval for AI apps: semantic + BM25 + temporal + graph recall over Qdrant/Memgraph, RRF fusion, 8-stage ranking, multi-tenant filters, answer synthesis.
mnemostack MCP server demonstrates solid definition quality with comprehensive parameter schemas and descriptions across 8 well-scoped tools. Most tools follow verb_noun naming conventions and have structured input/output documentation. However, there are gaps: descriptions lack LLM optimization guidance (many exceed 200 chars and bury key details), output schemas are not explicitly documented in the visible code, and error handling patterns are absent from the tool definitions. The server benefits from fastmcp framework which enforces Pydantic schemas, but the schema quality varies, some parameters like 'timestamp', 'offset', and 'sources' could benefit from format/range documentation. Security posture is reasonable (no exposed secrets in params visible), but there's no evidence of permission gates or audit trail documentation for destructive operations (mnemostack_remember, mnemostack_invalidate, mnemostack_feedback, mnemostack_graph_triples). The tools themselves are well-composed and serve a coherent memory/recall domain, with clear separation of concerns (remember vs search vs answer vs graph operations).
Generate an answer by recalling relevant memories and synthesizing them with an LLM.
Record user feedback signals (useful/irrelevant/clicked) for a recalled result to improve future ranking.
Execute a Cypher query against the Memgraph knowledge graph (only available if memgraph_uri is configured).
Ingest knowledge graph triples (subject, predicate, object) into Memgraph.
Check the health of the mnemostack server and its dependencies (Qdrant, Memgraph, embedding/LLM providers).
Mark memories as stale/invalid (soft delete). Invalidated facts are hidden from recall by default but accessible with include_invalidated=true.
Output schemas not documented in tool definitions. Visible code shows input schemas via Pydantic but does not declare what fields/types are returned. LLMs cannot plan downstream tool calls or extract data without knowing response structure.
Descriptions are verbose and lack LLM-optimization. 'Recall memories via semantic, BM25, and temporal retrieval with RRF fusion and optional 8-stage ranking pipeline' (106 chars, acceptable) but many include implementation details (RRF, ranking pipeline) rather than use-case clarity. Should answer: what does it do, when to call it, what does it return? Example: mnemostack_search description buries the key action.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 64 | 2026-07-28+ | v2 |
Ingest and embed memories into the collection. Accepts plain text (up to REMOTE_MAX_TEXT_CHARS) or chunked documents. Returns the ingested point IDs.
Recall memories via semantic, BM25, and temporal retrieval with RRF fusion and optional 8-stage ranking pipeline.
No error handling patterns or recovery guidance documented in tool descriptions. Destructive operations like mnemostack_invalidate and mnemostack_remember do not state what happens on failure, whether operations are retryable, or how to recover. Missing: error classification (retryable vs user-fixable vs fatal).
Parameter constraints underspecified. 'timestamp' (ISO-8601) and 'offset' (position within source) lack explicit range or format hints in descriptions. 'limit' (1-100) is constrained but could be clearer. 'token_budget' is unbounded in description.
No permission gates or scope declarations visible. Tools like mnemostack_invalidate (soft delete), mnemostack_remember (write), and mnemostack_graph_triples (write) do not declare required permissions (e.g., 'write:memory', 'admin:graph'). This prevents least-privilege agent configuration.
Tools accept complex nested objects (filters, metadata) without schema details. The 'metadata' param in mnemostack_remember is 'object' type with description 'Free payload fields, filterable at recall', LLMs have no guidance on valid field names, types, or nesting depth. Similarly, 'filters' in search/answer tools lack structure documentation.
mnemostack_graph_query tool description is minimal ('Execute a Cypher query against the Memgraph knowledge graph (only available if memgraph_uri is configured)') and does not guide LLMs on when to use it vs semantic recall. No examples of Cypher patterns, expected result structure, or error cases.