MCP server that gives AI agents persistent memory of your codebase. Remembers architecture decisions, patterns, conventions, and project context across sessions.
The server provides well-defined tools with consistent naming conventions and generally clear input schemas. All five tools follow verb-noun patterns (remember, recall, update_memory, forget, project_summary). Input schemas are present and properly typed with descriptions for most parameters. However, there are notable gaps in output schema documentation, the server returns JSON text responses but does not formally document the response structure for LLM consumption. Additionally, error handling lacks recovery guidance, and some parameter descriptions could be more actionable. The tools are well-composed and address a coherent domain (codebase memory), but definition quality is held back by incomplete response documentation and missing error classification patterns.
Delete a memory by ID. Use when information is outdated or incorrect.
Get an overview of all stored memories — counts by category, top tags, recent updates, and high-importance entries. Call this at the start of a session to load context.
Search and retrieve memories. Use natural language queries, filter by category, tags, or file paths. Returns the most relevant memories for the current task.
Store a new memory about the codebase — architecture decisions, patterns, conventions, bugs, or context. The AI agent calls this to save knowledge for future sessions.
Update an existing memory entry. Provide the memory ID and fields to change.
Output schemas not formally documented. Tools return JSON text via CallToolResult but do not declare the expected response structure (field names, types, cardinality). LLMs cannot reliably parse or chain results without explicit schema documentation.
Error handling lacks recovery guidance. When update_memory or forget fails, responses return bare error messages (e.g., 'Memory not found') without suggesting next steps (e.g., 'Use recall() to find the correct memory ID'). This forces LLMs to guess remediation.
recall() accepts optional query but does not document precedence. When multiple filters (query, category, tags, filePath) are provided, behavior is unspecified. LLMs cannot predict whether filters AND or OR, or which takes precedence.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | B | 70 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 51 | - | v1 |
project_summary description does not specify output structure or when to call it. The description says 'Get an overview... counts by category, top tags, recent updates, and high-importance entries' but does not explicitly state these are the fields returned, and omits a discovery/initialization hint.
Default importance value (5) is stated in description but not enforced in schema. If an LLM omits importance from remember(), the database assigns a default, but the schema does not declare this formally, risking silent inconsistency.