Outcome-based persistent memory for AI coding tools (Claude Code, OpenCode, Cursor)
Roampal's MCP server (inspect_server.py) defines 6 tools with basic structure but significant quality gaps. All tools have names following verb_noun convention (search_, add_, update_, delete_, score_, record_) and explicit input schemas with proper JSON Schema structure. However, parameter descriptions are universally minimal or absent, and output schemas are completely undocumented. The 'inspect' nature of this server (mock implementation returning static responses) limits real-world utility, but the quality issues would persist in production. Average tool quality across the 6 tools: 38/100.
Store permanent facts about the user
Delete a memory by ID
Store key takeaways from significant exchanges. NOT for permanent preferences or standing rules — use add_to_memory_bank for those.
Score previous exchange outcomes
Search across memory collections for relevant context
Update an existing memory
No output schemas documented for any tool. The call_tool handler returns a mock TextContent response, but the actual backend response structure is undefined. LLMs cannot plan downstream calls or extract typed fields without knowing what each tool returns.
Parameter descriptions are minimal or missing detail. 'memory_id' lacks context (is it a UUID? a string key?). 'memory_scores' is an object type with no description of expected structure or valid keys. 'key_takeaway' lacks format guidance. These gaps force LLMs to guess at valid input shapes.
Tool descriptions lack context for when to use them. 'search_memory' has no explanation of search scope, ranking, or result limits. 'score_memories' description ('Score previous exchange outcomes') is vague, what does 'scoring' mean? How is the score used? This ambiguity leads to incorrect tool selection.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 58 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 31 | - | v1 |
No error handling guidance. If search_memory returns no results, what does the agent do? If delete_memory fails (e.g. ID not found), is there a recovery path? The mock call_tool handler ignores the tool name entirely and returns a generic message. No actionable error messages for invalid input.
Destructive operations (delete_memory) lack confirmation/dry-run support. No pattern for preventing accidental deletion. No idempotent guarantees documented. Agents cannot safely retry failed delete calls.
No pagination or result limits documented. If search_memory returns hundreds of matches, there is no guidance on max results, offset/limit parameters, or how to iterate. Large result sets will exhaust token budgets.
Parameter naming lacks specificity in some cases. 'memory_id' is clear, but 'memory_scores' as an object type is opaque, what are the valid keys? Are they memory IDs? How should scores be formatted (0-1, 1-10, -1 to +1)? Requires documentation or an enum schema.
Tool composition gap: no documented relationship between tools. If update_memory requires a valid memory_id from search_memory, this dependency should be documented. If add_to_memory_bank returns an ID, it should be stated so agents can immediately update or reference it.