This server has 13 tools with generally clear naming and adequate descriptions, but falls short of production-grade quality on several fronts. Tool names follow verb_noun conventions (create_memory, search_memory, delete_memory, etc.), which is excellent. However, output schemas are not documented anywhere in the provided code, only input schemas are visible. Parameter descriptions exist but lack critical constraint details (format, range, valid values). Error handling is minimal: no evidence of recovery guidance, error classification, or actionable error messages. The server uses STDIO transport, which is a hard cap at 50 for protocol readiness. Most tools score in the 60-70 range individually; the average reflects decent naming and schema presence but gaps in descriptions completeness and missing output documentation.
Tools (13)
create_memorywritesource verified80/100
Store a new memory with semantic embedding for later retrieval
create_memory_relationwritesource verified77/100
Create a relationship between two memories (parent-child, related, etc.)
Output schemas are not documented. Only input schemas are visible; no documentation of what fields or structure the agent should expect in responses. This forces LLMs to guess at response structure and increases hallucination risk.
Parameter descriptions lack constraint details. For example, 'threshold' in search_memory is described as 'Similarity threshold (0-1, default: 0.7)' but does not state whether 0 and 1 are inclusive, or what happens if an out-of-range value is passed. 'limit' and 'offset' parameters lack minimum/maximum bounds. 'metadataFilter' in list_memories is documented as 'JSON string of metadata filters' but does not specify the structure or valid keys.
search_memorylist_memoriesget_memory_graph
Recommendations
Document output schemas for all 13 tools. For example, search_memory should document that it returns an array of objects with fields: id (string), content (string), similarity (number), metadata (JSON object). Use a JSON Schema format or clear prose describing the response structure.
Add constraint details to all numeric and enum parameters. For threshold, state: 'Range: 0.0 - 1.0 (inclusive). Memories with similarity below this value are excluded.' For limit and offset, state: 'Range: 1 - 1000. Default 50. Requests exceeding the maximum will be capped.'
Add confirmation or dry-run support for destructive tools. Example: add a 'confirmDelete' boolean parameter (as already done for delete_project) to delete_memory and delete_memory_relation. Or document a 'dry_run' mode that returns the would-be impact without executing.
Enhance discovery tool descriptions. For list_projects, change description to: 'List all available project namespaces. Call this first to see which projects exist before switching or creating memories.' For get_current_project, add: 'Returns the currently active project context. Useful to verify scope before memory operations.'
Define metadata schema explicitly. Document: 'metadata must be a valid JSON object with string keys and string/number/boolean values. No nested objects. Maximum 10 keys. Reserved keys (reserved_key_1, reserved_key_2) are forbidden.' If metadataFilter supports filtering by keys, provide example syntax: metadataFilter='{ "category": "important", "user_id": "123" }'.
Score history
Overall score trend
↑ 1 points across a rubric change (v1 → v2)
59/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-21
D
59
2026-07-28+
v2
2026-03-09
D
58
-
v1
read onlysource verified73/100
Get all relationships for a specific memory
list_memoriesread onlysource verified78/100
List all memories with optional pagination and sorting
list_projectsread onlysource verified70/100
List all available projects/namespaces
search_memoryread onlysource verified80/100
Search through stored memories using semantic similarity
switch_projectwritesource verified72/100
Switch to a different project/namespace for memory storage
No error handling guidance visible. The code shows tool registration but no evidence of try-catch blocks, validation, or error categorization. Destructive tools (delete_memory, delete_project, delete_memory_relation) lack confirmation or dry-run patterns. LLMs will have no guidance on retryability or recovery.
Sparse parameter descriptions on discovery tools. list_projects, get_current_project, and list_memories have minimal guidance on when to call them or what data they expose. Description text should include 'Call this first to discover available projects' or similar dependency hints.
Metadata parameters accept 'JSON string' but lack schema specification. 'metadata' in create_memory and update_memory, and 'metadataFilter' in list_memories are documented as strings but the actual JSON structure (keys, types) is not defined. This forces LLMs to guess or invent structures.
Ambiguity in tool naming for relations. 'get_memory_relations' and 'get_memory_graph' are similar, both retrieve relationships. The distinction (get_memory_relations returns direct relationships; get_memory_graph builds a traversed graph) is not clear from names alone. Descriptions should make this explicit to avoid LLM confusion.
No documented pagination or result limits for list_memories and get_memory_graph. list_memories defaults to limit=50 but no maximum is documented. get_memory_graph defaults depth=2 but no maximum is specified. Large result sets can blow the context window; limits should be documented and enforced.
list_memoriesget_memory_graph
Clarify the distinction between get_memory_relations and get_memory_graph in names or descriptions. Consider renaming one, e.g. 'get_direct_memory_relations' vs 'get_memory_graph' to make the scope clear. Or update descriptions: 'get_memory_relations: Returns one-hop relationships (parents and children). get_memory_graph: Builds a multi-hop graph up to N levels deep.'
Document result limits explicitly. For list_memories, state: 'Returns up to limit memories (max 1000 per request). Use offset for pagination. If total memories exceed 1000, use multiple calls with different offsets.' For get_memory_graph, state: 'Maximum depth: 3. Graphs exceeding 500 nodes will truncate with a truncated=true flag in the response.'
Add error message examples in descriptions. For example, update_memory could state: 'On error: returns {error: "Memory not found", id: "..."}. Try search_memory to find the memory ID.' Or delete_project could state: 'On error: returns {error: "Project not empty", memories_count: 42}. Delete all memories in the project first, or set confirm_cascade=true to delete all.'
Implement input validation and return actionable errors. Ensure tools validate relationType enum values, threshold range, limit bounds, etc., and return clear messages: 'Invalid relationType: got "weird". Must be one of: parent-child, related, follows-from, contradicts, updates, supports.' This guide LLM self-correction.
Add field consistency across related tools. Ensure search_memory, list_memories, and get_memory_graph all return the same memory object structure (same id, content, metadata field names). This enables tool chaining without the LLM having to remap field names.