An MCP server that implements a persistent memory layer for AI agents, storing and recalling architectural decisions, bug fixes, and coding patterns from Git repositories using semantic search and hybrid ranking.
The server defines 4 tools with explicit schemas and descriptions visible in src/mcp_server.py. Tool names use action verbs (recall, store, list, flag) which is good naming practice. However, descriptions are inconsistent in depth, recall_memory and store_memory have detailed, context-aware descriptions (120-180 chars), while list_recent_memories and flag_contradiction are shorter and less prescriptive (40-80 chars). All tools have input schemas with typed parameters and defaults, but output schemas are not formally documented, responses are unstructured strings rather than typed objects. Error handling exists but is minimal: catch-all Exception blocks return generic error messages without recovery guidance. Parameter descriptions are present but inconsistent, some lack detail about format or constraints. No tool annotations (readOnlyHint, destructiveHint, idempotentHint) despite clear read/write semantics. This is typical mid-tier community work: functional but not production-grade.
Let the agent report a conflict or state that a memory is obsolete. This lowers the confidence of the memory and tags it.
Surface what is known about a specific module sorted by recency.
Search the AI Memory Layer for past architectural and coding decisions. Use this tool when trying to understand why a certain technical choice was made, what conventions exist in the codebase, or how past bugs were resolved.
Actively save a new architectural decision, bug fix, or pattern mid-conversation. memory_type must be one of: episodic, semantic, procedural.
No documented output schemas. All tools return unstructured strings (e.g., 'No relevant memories found.' vs formatted memory objects). LLMs cannot plan downstream tool calls or extract structured data without knowing response fields.
Missing tool annotations. store_memory and flag_contradiction are destructive writes but carry no destructiveHint. recall_memory and list_recent_memories are read-only but lack readOnlyHint. This prevents safety-aware clients from enforcing operation classes.
Weak error messages. All exceptions caught and returned as strings like 'Failed to store memory: {e}' without recovery guidance. LLMs cannot determine if an error is retryable, user-fixable, or fatal.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 66 | 2026-07-28+ | v2 |
Parameter validation is lax. memory_type in store_memory states 'must be one of: episodic, semantic, procedural' but no enum constraint in schema. module parameter is optional/nullable with no validation. LLMs can pass invalid values.
Inconsistent parameter descriptions. limit parameter in recall_memory (72 chars) and list_recent_memories (73 chars) lacks guidance on reasonable bounds. No mention of pagination strategy or performance implications if limit is high.
Missing idempotency guarantee documentation. store_memory checks for duplicate content_hash, making it idempotent, but this is not documented in the description. Agents need to know whether retry is safe.