A platform-agnostic persistent memory layer for AI agents
Three tools with solid naming (verb_noun pattern) and reasonable descriptions (94-150 chars), but significant gaps in schema completeness, parameter validation, and error handling. Tool schemas are present and typed, but lack critical constraints like enums for state transitions, range limits on numeric parameters, and mutual exclusivity documentation. Error handling is minimal, the forget() tool returns a bare error object without recovery guidance. Output schemas are implied but not formally documented. The codebase shows thoughtful domain design (Buddhi evaluation, scope hierarchies) but lacks production-grade validation and LLM-optimization patterns.
Transition memories to latent or dissolved state. Latent memories are dormant but recoverable. Dissolved memories leave only a trace record. Target by specific memory ID or by scope subtree.
Retrieve relevant memories for a given context. Performs semantic search across all stored memories and returns composite-scored results blending similarity, importance, and recency.
Store a memory through the Antaḥkaraṇa pipeline. Buddhi evaluates the content for importance, scope, and categories before storing in Chitta. Use this to persist knowledge, decisions, preferences, or any information worth remembering across sessions.
forget() tool lacks mutual exclusivity documentation and recovery guidance. Parameters memory_id and scope are mutually exclusive, but the function body returns a bare error dict without instructing the LLM what to do next. Error response: '{"affected": 0, "error": "Must specify either memory_id or scope."}', LLM cannot infer whether to retry, ask user, or call a different tool.
remember() tool accepts importance as unconstrained float (no min/max specified). Description says '0.0 to 1.0' in English but JSON Schema has no type bounds. An LLM could pass importance=2.5 or importance=-1.0, causing silent failure or incorrect Buddhi logic. Spec says: 'Specify minimum and maximum for numeric parameters.'
forget() new_state transition ('latent' vs 'dissolved') is not exposed as a user-facing enum or clear parameter. The tool logic hardcodes state transitions based on force_dissolve flag, but the description does not explain what state values are valid, when transitions fail, or what states are allowed from each prior state. This violates the pattern:constrained-input principle that enums prevent hallucinated values.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 55 | <=2025-11-25 | v2 |
| 2026-03-09 | D | 55 | - | v1 |
Output schemas for all three tools are undocumented. The source shows tools return JSON-serialized dicts (e.g., remember() returns {"stored": bool, "memory_id": str, ...}), but the MCP tool definitions do not formally declare the output schema structure. LLMs cannot plan downstream calls or validate responses without knowing the field types and presence guarantees.
recall() tool has no pagination or result limiting enforcement. The 'limit' parameter defaults to 5, but if Buddhi evaluates 100s of memories, the search could return all of them, exploding token usage. No mention of total_count or next_cursor in the response schema, and no guidance on handling large result sets.
Parameter descriptions lack format/constraint details. E.g., scope parameter is described as 'Override Buddhi's scope inference (e.g. /project/jozu/architecture)', the example suggests a path format, but no pattern, character restrictions, or depth limits are specified. An LLM might pass scope='/...invalid...path' or scope='', causing silent failures.
No input validation or actionable error messages. If scope is invalid, if importance is out of range, or if Buddhi API fails, the tools do not return clear recovery guidance. E.g., remember() returns {"stored": False, "reason": "Buddhi determined..."}, but the LLM has no way to know whether to retry, ask the user, or call a different tool.
source_agent parameter is optional and undocumented. The description says 'Identify which agent is storing this (claude-code, openclaw, etc.)', these examples might mislead the LLM into thinking only those specific values are valid. Should define an open string or provide a clear enum of valid agent names.