Persistent memory system for coding agents - exposes memory as first-class tools to any MCP-compatible host
RECALL presents a well-intentioned memory system for agents with 15 tools spanning search, retrieval, and recording functions. However, the evaluation reveals critical gaps in parameter descriptions, missing output schemas, and insufficient error handling guidance. While tool names are action-oriented (memory_hybrid_search, add_breadcrumb, get_recent_decisions) and follow verb_noun convention, the rubric's hard constraints apply: many parameters lack descriptions or detail, and output structures are not documented. The server uses STDIO transport only, which triggers a hard cap of 50 on protocol readiness. The codebase shows structured implementation but does not expose complete schema documentation in accessible form. Most tools fall in the 50-65 range individually; averaging to 58 overall.
Add a breadcrumb (context, note, reference) to memory
Record a decision to memory
Record a learning (problem + solution) to memory
Retrieve context for an agent by running hybrid search and formatting results
Find decisions similar to a given decision
Retrieve a specific decision by ID
Retrieve a specific LoA (Lesson of Arms) entry by ID
Retrieve recent breadcrumbs
No output schemas documented for any tool. Tools do not declare what fields they return, preventing agents from planning downstream operations or extracting needed data for chaining.
Parameter descriptions are uniformly terse (5 - 20 characters). 'Optional project filter' and 'Maximum results to return' lack format guidance, valid ranges, or dependencies. Rubric hard constraint: descriptions under 20 chars cannot exceed 40/100.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | 2026-07-28+ | v2 |
Retrieve recent active decisions
Retrieve recent learnings
Retrieve recent LoA (Lesson of Arms) entries
Get statistics about memory database contents
Hybrid search combining FTS5 + vector embeddings with RRF fusion
Mark a decision as reverted (was wrong, rolled back)
Mark a decision as superseded by a newer decision
Importance and confidence parameters lack type and range constraints. No enums, no min/max. 'importance': 1-10 is mentioned only in input description; agents will guess valid ranges and may pass invalid values.
No error handling guidance. Tools do not document what happens on invalid inputs, duplicate entries, missing IDs, or constraint violations. Agents have no recovery path.
Domain jargon ('LoA', 'Lesson of Arms') used in tool names and descriptions without explanation. Agents unfamiliar with this term will struggle to recognize when to call these tools.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint). Agents cannot distinguish safe (read-only) operations from risky ones (add_decision, supersede_decision). This matters for retry safety and planning.
Pagination not evident. get_recent_breadcrumbs, get_recent_decisions, etc. accept a 'limit' parameter but do not declare total count or next_cursor. Large result sets risk blowing context windows without pagination guidance.
Duplicate/overlapping tools. get_recent_decisions, get_decision, and find_similar_decisions serve slightly different purposes but lack clarity on when to call each. Agents waste reasoning cycles disambiguating.