MCP server giving Claude persistent autistic-style memory + learning
The server presents 15 tools with significant structural quality variation. Tools like memory_recall and memory_capture show detailed schema definitions with proper parameter typing, good descriptions (120-350 chars), and well-documented input constraints. However, 8 of 15 tools (53%) have minimal or absent descriptions (<50 chars) and no visible input schemas in the provided source, forcing conservative scoring. The codebase references 'mcp-wrapper/src/tools.ts' where definitions should live, but that file is not included in the source code sample provided. Memory management tools (recall, capture, reinforce, contradict, consolidate, search, structural) show good intent but inconsistent delivery. Profile, schema, topology, events, episodes, claim_check, and temporal_recall tools lack visible schemas and descriptions, these appear to be inferred from tool names only. Error handling is absent across all tools; no recovery guidance is provided. The server does implement tool annotations (readOnlyHint/destructiveHint visible for risk classification), which is positive, but this is insufficient to overcome schema and description gaps.
Verify or check claims against stored memories
Query pending curiosity/learning items
Retrieve recent episodes from the memory store
Query events in the memory store
Capture a verbatim turn (auto-dedups near-duplicates). Use for corrections, not for minting standing-order directives.
Consolidate memories during sleep cycles
Mark a record contradicted; new fact stored as a NEW record (old NEVER deleted). Mutates store.
8 of 15 tools (53%) have no visible input schemas. Tools like memory_search, memory_consolidate, profile_get_set, curiosity_pending, schema_list, events_query, topology, episodes_recent, memory_temporal_recall, and claim_check reference only names and risk classifications, no parameter definitions visible in provided source.
Tool descriptions for 8 tools are trivial or missing (under 50 characters). 'Get or set user profile information', 'Query pending curiosity/learning items', 'List available memory schemas', 'Query events in the memory store', 'Query graph topology of memories', 'Retrieve recent episodes from the memory store', 'Recall memories using temporal ordering', 'Verify or check claims against stored memories', all lack detail on WHEN to use, WHAT side effects occur, or WHAT is returned. This violates pattern:tool-description baseline (194 chars average for A+ tools).
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | F | 49 | 2025-06-18+ | v2 |
Recall verbatim memories by cue — decisions, preferences, prior discussion, rationale. Call before a repository search. Returns hits + anti_hits.
Recall memories using structural graph traversal
Boost Hebbian edges among co-retrieved record ids. Mutates edge weights. Use when two records co-answered.
Search the memory store with semantic or lexical queries
Recall memories using temporal ordering
Get or set user profile information
List available memory schemas
Query graph topology of memories
No input validation, constraints, or error guidance visible for any tool. Tools like memory_recall accept budget_tokens (integer), cue_embedding (384 floats), and language (8 supported ISO codes), but no error messages, no recovery guidance, no clear validation rules. All 15 tools lack this.
Output schemas are not documented for any tool. The rubric requires: 'Document the output schema. LLMs need to know what fields to expect.' Tools like memory_recall reference 'hits + anti_hits' but no structured return schema is visible. memory_capture, memory_reinforce, memory_contradict show partial input schemas but no output specification.
Tool names like 'memory_recall_structural' vs 'memory_recall', 'profile_get_set', and 'episodes_recent' lack clear action verbs or contain compound responsibilities. 'profile_get_set' combines GET and SET into one tool (violates pattern:tool, 'each tool should do exactly one thing'). 'episodes_recent' is a noun-verb hybrid (should be 'list_episodes' or 'get_recent_episodes'). Names do not clearly convey side effects, 'memory_capture' vs 'record_memory' is ambiguous on write intent.
Composition risk: memory_recall, memory_search, and memory_recall_structural appear to do similar things ('retrieve memories using different methods'). Tool chain is unclear, LLMs may waste reasoning cycles choosing between recall methods. No guidance on which to use when. Violates pattern:tool, 'Avoid multiple tools that do the same thing differently.'
Idempotency not documented. memory_capture mentions 'auto-dedups near-duplicates' but does not state whether calling twice with the same text is safe or produces duplicates. memory_reinforce says 'identical pair sets are idempotent within one session' but not across sessions. No clear guidance for LLM retry safety. Violates pattern:idempotent-operation.