Hosted persistent memory for AI agents that learns which facts help — feedback re-ranks recall — one key across Claude Code, Cursor, VS Code & ChatGPT, no infra to run
The server provides 10 well-named memory tools with structured schemas and reasonable descriptions. Tool names follow verb_noun convention (memory_read, memory_write, memory_search, etc.), which is clear and consistent. However, several tools lack sufficient detail in parameter descriptions, and output schemas are not documented in the source code provided. The schema definitions visible in the input show basic type information (string, number, array) but lack constraints like enums, min/max bounds, or validation rules. Error handling patterns are present (ApiError, NetworkError types referenced) but guidance text is not visible in the tool definitions themselves. The result truncation logic (capResult function) shows attention to token limits and output sizing, which is a positive signal. Overall, this is a solid B-/C+ tier implementation: functional and well-structured, but missing some production-grade polish around parameter constraints, output documentation, and error recovery guidance.
Create a new shared memory room
Get statistics about memory domains
Provide feedback to re-rank memory recall accuracy
Create an invitation to share a memory room
Read facts from memory matching a query
Retrieve recently added or modified memories
List available memory rooms (shared memory spaces)
Output schemas not documented. The source code snippet shows input schemas for all 10 tools, but no documented return types or response structures are visible. LLMs need to know what fields to expect (e.g., does memory_read return an array of {id, content, score}?). This prevents agents from planning downstream operations and forces them to infer structure.
Parameter constraints missing. The 'top_k' parameter (number) in memory_read and memory_search lacks min/max bounds. The 'limit' parameter in memory_recent has no range constraint. Without explicit bounds, LLMs may pass absurd values (e.g., top_k=999999) that break the API or timeout. The schema shows type but not range.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | D | 55 | 2025-06-18+ | v2 |
Search memory with advanced filtering and ranking
Access the memory vault with archived and pinned items
Write facts to memory with optional concepts
Vague parameter descriptions. 'outcome' in memory_feedback is described as 'Feedback outcome score (positive/negative)' but does not clarify the numeric range (is it -1/+1? 0-10? -5 to +5?). 'expires_in_days' and 'max_uses' in memory_invite_create lack minimum/maximum guidance. Ambiguous descriptions force LLMs to guess or require clarification loops.
No error recovery guidance visible in tool definitions. The source code references ApiError, NetworkError, and UnreadableBodyError types, and the comment notes 'all three are exported from ./shared', but the actual error handling logic and recovery messages (e.g., 'Use a more specific query') are defined elsewhere (teaching.ts, scope.ts). Tool definitions should embed actionable error hints so LLMs know what to do when a call fails.
Empty input schema for read-only discovery tools. 'memory_rooms_list', 'memory_domains_stats', and 'memory_vault' accept {} (no parameters). While this is valid, the descriptions do not explain what triggers these operations or when to call them. Discovery tools should hint at their purpose: 'Call this first to understand available rooms before creating one.'
Overlapping tool names and responsibilities. 'memory_read', 'memory_search', and 'memory_recent' all retrieve facts from memory but with different filters/orderings. The descriptions distinguish them ('Search with advanced filtering' vs 'Read matching a query' vs 'Recently added'), but the differences are subtle. An LLM may struggle to pick the right one. Consider whether these should merge into a single parameterized 'memory_query' tool, or clarify in descriptions when to prefer each.