Persistent memory for AI coding — semantic search, git history analysis, and intelligent context preservation
code-memory demonstrates solid foundational tool design with clear action verbs and reasonable descriptions, but has critical gaps in parameter documentation, schema detail, and error handling guidance. All 7 tools are explicitly registered with schemas, descriptions, and input parameters visible in src/mcp/tools.rs. Naming follows verb_noun conventions (search_code, explain_code, trace_decision, find_related, remember, index_project, get_session_patterns). Descriptions range 78-184 characters, meeting the 10-1024 baseline. However, several tools lack completeness in parameter typing, output schema documentation, and error recovery guidance. The remember and index_project tools lack structured output specification. No tools include error classification or recovery guidance. Tool composition is reasonable, each has a single clear purpose, but several parameters are under-documented regarding constraints, valid ranges, and mutual exclusivity.
Get detailed explanation of a code symbol — its definition, usage, dependencies, and git history.
Find code related to a given file via dependency graph traversal. Shows what depends on this file and what it depends on.
Retrieve learned patterns from Claude Code session history. Shows coding conventions, error-fix pairs, architectural decisions, and other patterns discovered from past sessions.
Manually trigger reindexing of the project. Use this when files have changed and you want to refresh the search index.
Store a piece of knowledge for future recall — architectural decisions, patterns, conventions, or any context that should persist across sessions.
Search code using full-text and semantic search. Finds relevant code snippets by keyword or natural language meaning.
Output schemas not documented for any tool. The rubric requires 'Document the output schema. LLMs need to know what fields to expect so they can plan downstream tool calls and extract the right data.' Without documented return types, agents cannot reliably chain tools or extract results.
No error handling or recovery guidance in any tool description. The rubric requires 'Error responses must tell the LLM what to do next' and 'Categorize errors as retryable, user-fixable, or fatal.' Currently, tool descriptions do not indicate what errors are possible, how to recover, or what the LLM should do if a call fails.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | D | 56 | <=2025-11-25 | v2 |
Trace architectural decisions from git history — find why code exists, who changed it, and the rationale behind changes.
Parameters lack constraint documentation. 'max_results' accepts integer with default 10, but no minimum/maximum bounds stated (baseline requires 'Specify minimum and maximum for numeric parameters'). 'depth' in find_related similarly unbounded. 'min_confidence' in get_session_patterns states range (0.0-1.0) but others lack this clarity.
'time_range' parameter in trace_decision documents valid values ('7d', '30d', '90d', '1y', 'all') in description but should be formalized as an enum in the JSON schema for machine-parsability. Same issue with 'direction' in find_related, which IS correctly enumerated, inconsistent schema quality.
remember tool (WRITE risk) lacks confirmation/dry-run pattern. The rubric states 'Irreversible operations (delete, send, publish) should support a dry-run or confirmation step.' Storing knowledge that could override prior context is a state-modifying operation but offers no preview or confirmation mechanism.
index_project tool (WRITE risk) description is too brief (65 chars, below baseline 194 avg). 'Manually trigger reindexing of the project. Use this when files have changed and you want to refresh the search index.' Does not state what happens to existing index, estimated duration, or failure modes. LLMs cannot reason about when/if to call it.
Parameter relationships not documented. 'path' parameter in search_code interacts with 'language' filter, but no description states whether both can be applied together, which takes precedence, or how they combine. 'include_git_history' in search_code similarly lacks documentation on what structure the git history response takes.
get_session_patterns lacks clarity on what 'patterns' are returned. Description says 'Retrieve learned patterns from Claude Code session history. Shows coding conventions, error-fix pairs, architectural decisions...' but does not define the output structure, fields, or how to use returned patterns in subsequent tools.