Compact, efficient, and extensible long-term memory system for LLM agents with structured semantic memory, procedural skill memory, and session management
LycheeMem presents a specialized memory-augmentation tool suite with clear domain focus. Tools are well-named with action verbs (smart_search, append_turn, consolidate) and have detailed descriptions exceeding 150 characters. However, significant gaps exist: (1) parameter descriptions are partially in Chinese, reducing English-language LLM clarity; (2) output schemas are not documented, we cannot see what fields responses contain; (3) error handling guidance is absent; (4) tool annotations (readOnlyHint, destructiveHint) are missing despite write operations present. The Hermes plugin variant (tools.py) shows marginally better English descriptions but adds enum constraints and validation guidance absent from the base MCP schema. The HTTP transport is modern, but lack of response schema documentation violates pattern:tool-chain since downstream tools cannot plan calls without knowing what fields are returned.
Append one natural-language conversation turn from an external host into LycheeMem's session store. Call it after every completed dialogue turn so both the user turn and the assistant reply are mirrored into the same session_id, even if you do not consolidate on that turn. Do not use it for raw tool invocations, tool arguments, tool outputs, or other orchestration-only traces unless the host explicitly wants those artifacts stored as memory.
Persist new long-term memory after a conversation. Call it only after the relevant natural-language user and assistant turns have already been mirrored with lychee_memory_append_turn. Use it when the conversation introduced new facts, entities, preferences, relationships, or reusable procedures that should be stored for future retrieval.
Developer-facing raw retrieval tool. Retrieve relevant information from LycheeMem structured long-term memory when you explicitly want the raw semantic_results (legacy alias: graph_results) and skill_results payload. For normal agent use, prefer lychee_memory_smart_search.
Primary recall path for agents. Performs LycheeMem search and can return a compact background_context built directly from retrieved memory. Prefer response_level=minimal for normal agent use, and switch to full only for debugging.
No output schemas documented for any tool. LLMs cannot infer what fields responses contain (e.g., what does background_context include? Is there a total count? Pagination cursors?). This violates pattern:tool-chain, downstream tools cannot be selected without knowing what IDs and references are returned.
Missing tool annotations (readOnlyHint, destructiveHint) for write operations. lychee_memory_append_turn and lychee_memory_consolidate modify state but carry no destructiveHint flag. Agents cannot assess risk, they should warn before or require confirmation for state-modifying calls.
Parameter descriptions inconsistently in Chinese (lychee_memory_smart_search, lychee_memory_search) vs English. English-language LLMs rely on English descriptions for parameter interpretation. Chinese descriptions are opaque and fail LLM parameter binding.
Inferred effective spec: 2025-06-18+.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 69 | 2025-06-18+ | v2 |
Enum-like parameters not constrained as enums in JSON Schema. 'mode' (raw/full/compact) and 'response_level' (minimal/compact/full) are described as having specific values, but the JSON Schema uses type:string without enum constraints. LLMs may hallucinate invalid values. Hermes variant correctly uses 'enum' constraint.
No error handling guidance. Tools lack descriptions of failure modes (e.g., session_id not found, semantic indexing errors, consolidation timeout). Agents cannot self-correct or recover, they receive raw errors with no actionable next steps.
Two nearly identical search tools (lychee_memory_search vs lychee_memory_smart_search) with almost identical inputs cause LLM confusion. Descriptions say one is 'primary' and one is 'developer-facing raw', but without output schema differences, LLMs cannot reliably distinguish when to choose which. Violates composition guideline: 'Avoid multiple tools that do the same thing differently.'
Default parameter values not justified. 'top_k: 5' and 'response_level: minimal' are set, but no explanation of why these defaults were chosen or what common use case they target. If an agent forgets top_k, it silently returns only 5 results instead of the needed 50, no warning.