MCP server for storing and retrieving episodic memories with semantic search capabilities, vector embeddings, and DuckDB-backed persistence
The Engram server provides 5 well-structured tools with comprehensive schemas and detailed descriptions. Tool naming follows verb-noun conventions (add_memory, search, get_episodes, update_episode, get_status). All tools have detailed parameter documentation with types and constraints. Output schemas are implicitly well-defined through DuckDB integration and semantic search parameters. However, some descriptions are verbose (search tool description exceeds 1024 char rubric guideline by ~50%), and error handling guidance is minimal. Tool composition is strong, each tool has a single responsibility. Security consideration: the server correctly uses source/group_id for multi-tenant isolation rather than secrets in parameters. The search tool documentation is exceptional in guiding mode selection (vector vs keyword vs hybrid), which matches the pattern of actionable parameter descriptions. No tool descriptions are missing, and all parameters have type annotations.
Store a new episode in memory
Retrieve recent episodes in chronological order. All parameters are optional — call with no arguments to get the most recent episodes.
Get the system status including database statistics, embedding model info, and health metrics.
Search episodes using semantic similarity, keyword matching, or hybrid mode. For most searches, only provide 'query'. All other parameters are optional secondary filters — omit them unless you have a specific reason to narrow results. Search mode guidance: - hybrid (recommended): best for most queries — balances semantic understanding with exact term matching. - vector: best for concept/intent queries where your words won't match the stored text (e.g. "deployment preferences" finding CI/CD memories). - keyword: best for exact terms, proper nouns, error codes, or version strings where semantic drift would hurt (e.g. "mlx_lm.server"). Note: the default search_mode will change from 'vector' to 'hybrid' in the next major version.
Update metadata, tags, or expiration of an episode. Setting expired_at to a past timestamp performs a soft-delete — the episode is hidden from default search but remains recoverable by setting expired_at back to null. Use tags to demote (e.g. add 'deprecated') so callers can filter stale content at query time.
search tool description exceeds recommended length (~1100 chars vs 1024 rubric max). While comprehensive, this wastes tokens and buries key guidance in prose.
No recovery guidance in error cases. LLM callers have no direction if add_memory fails with duplicate, update_episode targets non-existent ID, or search returns no results.
get_status has minimal description (24 chars), does not explain what statistics or metrics it reveals or when to call it for monitoring vs debugging.
Metadata and JSON string parameters (metadata field in add_memory, update_episode) lack schema definition. LLMs cannot infer valid structure for 'JSON string with additional metadata', should provide examples or enumerated keys.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | B | 77 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 47 | - | v1 |
No documented output schema for search results. Response likely returns vector similarity scores, timestamps, and content, but this is not explicit in tool definition, forcing LLMs to infer structure.