Zikkaron presents a sophisticated memory engine with 21 tools, but falls short of production-grade quality. All tools have descriptions and explicitly registered schemas (visible in server.py), but descriptions are often domain-specific and assume deep knowledge of neuroscience-inspired memory concepts. Parameter descriptions are present but frequently lack actionable constraints, enums, or format specifications. Output schemas are not documented in the source code provided. Error handling guidance is minimal. Naming follows verb-noun conventions consistently (remember, recall, forget, etc.), which is a strength, but many tools do not clearly communicate when they modify state or irreversible side effects. Security posture lacks explicit permission gates and audit logging despite operating a stateful memory engine.
Mark a memory as protected/anchored (critical facts, decisions, constraints)
Analyze memory coverage gaps in a project domain
Save a snapshot of memory state for potential rollback
Trigger immediate memory consolidation (background daemon normally does this)
Create a prospective memory trigger (future-oriented reminder)
Detect knowledge gaps and missing connections in memories
Zoom into a specific memory and explore its connected memories
Output schemas not documented. Tools return complex structures (memories with embeddings, knowledge graphs, hierarchical results) but source code provides no formal documentation of return types, field names, or data structures. LLMs cannot plan downstream tool chains without knowing what fields to expect.
Descriptions assume deep neuroscience/memory-science domain knowledge (e.g., 'spreading activation', 'reconsolidation', 'prospective memory engine', 'Hopfield networks'). LLMs trained on general web text may not reliably infer correct behavior from these descriptions. No plain-English fallback explaining when/why to call each tool.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 59 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 48 | - | v1 |
Mark a memory as deleted or reduce its heat (relevance)
Retrieve hot memories filtered by project directory
Generate an autobiographical narrative of project events
Install Claude Code hooks for auto-capture and context injection
Get detailed statistics on stored memories by type and heat
Navigate the knowledge graph using semantic paths
Rate a memory's usefulness or update its importance
Retrieve memories matching a query with vector search and spreading activation
Retrieve memories using fractal hierarchical search across abstraction levels
Store a new memory with embedding and optional file hash. context MUST be the actual working directory path (e.g., '/home/user/projects/myapp'), NOT a description. get_project_context() filters by directory path match — descriptive strings will make memories unfindable by project.
Restore memory state from a previous checkpoint
Initialize a project context with seed memories
Synchronize system instructions and configuration with memory state
Check a memory's consistency with the knowledge graph
No state-modifying tool declares dry-run or confirmation support. Tools like 'forget', 'restore', 'checkpoint', 'install_hooks' can irreversibly alter memory state or install system hooks, but LLMs have no way to preview effects or confirm before execution. High risk of accidental data loss or misconfiguration.
No visible error handling or recovery guidance in tool descriptions. Descriptions say what tools do, but not what errors they can raise, how LLMs should interpret them, or what remediation steps are available. E.g., 'remember' provides detailed instructions about 'context MUST be the actual working directory path', but no guidance on what happens if a bad path is provided or how to recover.
Parameter descriptions lack actionable constraints. 'rating' parameter in 'rate_memory' says '0-1 or -1 to 1', which is correct? No enum, no examples of valid values, no guidance on impact. 'hook_types' in 'install_hooks' lists types in description but not as enum constraint. LLMs must infer valid values from text, inviting hallucination.
No visible permission gates or audit logging. The server modifies persistent memory, installs system hooks, and manages state across sessions, but tool definitions provide no scoping information (e.g., 'requires memory:write', 'requires system:admin'). No indication that calls are logged for compliance or incident response.
Composition and chaining undefined. 'remember' stores memories; 'recall' retrieves them. But response fields of 'recall' are not documented, does it return 'memory_id', 'memory_ids', or 'id'? If 'forget' expects 'memory_id', the chain breaks unless field names match exactly. No per-tool response documentation visible.
Result limits not enforced or documented. 'recall', 'recall_hierarchical', 'navigate_memory', and discovery tools do not specify max result sizes. A poorly-tuned hierarchical search could return thousands of memories, consuming the entire context window. No pagination support visible.
Tool names for administrative tasks are unclear. 'consolidate_now', 'sync_instructions' do not convey urgency, scope, or risk. LLMs may call them casually without understanding they trigger background processing or synchronization. Renames like 'trigger_memory_consolidation' or 'refresh_system_instructions_from_memory' would be clearer.
'directory' parameter expected to be a 'working directory path' but no validation rules documented (must exist? must be absolute? symlinks allowed?). 'remember' emphasizes this in description, but 'get_project_context', 'seed_project', 'assess_coverage', 'detect_gaps', 'get_project_story' reuse 'directory' param without that same guidance, risking misuse.