A reflective memory system that remembers and searches documents by meaning. Provides semantic memory storage, vector search, and integration with LangChain agents and OpenClaw context engine. Includes MCP server for tool-based access to memory operations.
Four tools with adequate naming and descriptions, but significant schema and documentation gaps. Tool names follow verb_noun convention (remember, recall, get_context, update_context), which is positive. Descriptions range 79-138 chars, within the productive range but minimal for LLM disambiguation. Critical issue: input schemas are inferred from the toolkit.py file snippet and not fully visible in structured form; parameter descriptions exist but lack type constraints (enums, patterns, min/max bounds). The 'tags' parameter in remember() accepts an object with no schema structure provided. No output schemas are documented. Error handling strategies are absent. Two tools (remember, update_context) are write operations but lack explicit confirmation patterns or idempotency guarantees.
Read the current working context — active goals, recent decisions, and state. Check this at the start of a session or when you need to orient.
Search long-term memory for relevant past notes, facts, preferences, or decisions. Use natural language queries.
Store a fact, preference, decision, or important note in long-term memory for later recall. Use this when the user shares something worth remembering.
Update the current working context with new state, goals, or decisions. This persists across sessions.
Input schemas lack type constraints and detailed parameter descriptions. 'tags' parameter in remember() accepts untyped object with no schema validation. 'limit' in recall() has default but no min/max bounds. 'content' parameters in remember() and update_context() lack length limits or format guidance.
No output schemas documented. LLMs cannot plan downstream operations or extract structured data from recall() or get_context(). Unknown whether recall returns paginated results, how many fields are included, or what fields identify memories for later operations.
Write operations (remember, update_context) lack explicit state-change language in descriptions and provide no confirmation or dry-run pattern. Descriptions do not clearly state 'This persists data' or 'This is irreversible', critical for agent safety.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 62 | 2026-07-28+ | v2 |
| 2026-06-11 | F | 43 | - | v1 |
No error handling guidance. Descriptions lack recovery hints (e.g., 'If query returns no results, try simpler keywords' or 'Content may not persist if storage backend is unavailable'). No categorization of retryable vs. fatal errors.
Parameter descriptions are sparse and generic. 'tags' is described as 'Optional tags to categorize the memory. Example: {"topic": "preferences", "project": "myapp"}', the example value risks being reused literally by LLMs rather than adapted. Should use enum or pattern constraints instead.