Lightweight, embedded graph-based memory system for AI applications with MCP (Model Context Protocol) integration for Claude Desktop, Claude Code, and Cursor
KuzuMemory defines 6 tools with mostly present descriptions and schemas. Naming is clear and verb-forward (kuzu_enhance, kuzu_learn, kuzu_recall, kuzu_remember, kuzu_stats, project). Descriptions are detailed and context-rich (100-250 chars), well above the 10 - 1024 char baseline. Input schemas are present with proper JSON Schema structure. However, several critical gaps reduce the score: (1) Parameters lack explicit type definitions in some cases, 'prompt' and 'content' are strings but lack min/max length constraints; 'max_memories' and 'limit' are integers but lack range boundaries (should enforce e.g. 1 - 100 to prevent DoS). (2) Output schemas are NOT documented, the tool descriptions explain what is returned conceptually ('memories ranked by relevance', 'statistics'), but there is no formal JSON Schema definition of response shape. (3) Error handling is not visible in any tool definition; no recovery guidance ('If search fails, try...'), no categorization of retryable vs fatal errors. (4) No tool annotations (readOnlyHint, destructiveHint, idempotentHint) present despite clear semantic differences: kuzu_remember and kuzu_learn are writes, kuzu_enhance/kuzu_recall/kuzu_stats/project are reads. (5) kuzu_learn and kuzu_remember docs state 'async/sync' but no actual async timeout or error path is documented. (6) Parameter relationships (e.g., memory_type enum values in kuzu_remember) are documented in descriptions but not validated in schema. Overall definition quality is solid but lacks production-grade rigor in output contracts and error surfaces.
RAG prompt augmentation: Enhance prompts with project-specific context from KuzuMemory using semantic search and vector similarity. Performs context injection by retrieving relevant project memories, patterns, and learnings to augment the input prompt. Use this for context-aware AI responses that understand project history and domain knowledge.
ASYNC/BACKGROUND/NON-BLOCKING continuous learning: Store observations, insights, and learnings asynchronously during conversations without waiting for confirmation. Ideal for capturing context, patterns, and evolving understanding as they emerge. Returns immediately without blocking. When to use: Ongoing conversation learnings, observations, insights, context capture during development sessions. When NOT to use: Critical facts requiring immediate confirmation (use kuzu_remember instead for synchronous storage of important decisions, preferences, or facts that must be stored immediately).
Semantic memory retrieval: Query project memories using vector search and similarity matching. Performs semantic search across stored learnings, patterns, decisions, and context to find relevant information based on meaning rather than exact keyword matches. Returns memories ranked by relevance score and temporal decay weighting. Use this to retrieve project-specific knowledge, past decisions, learned patterns, or context from previous conversations.
SYNC/IMMEDIATE/BLOCKING critical fact storage: Store important decisions, preferences, or facts that must be confirmed immediately. This operation waits for database confirmation before returning, ensuring the memory is durably persisted. Use for critical information that requires immediate storage verification. When to use: Important decisions, user preferences, project constraints, architectural choices, critical facts that need immediate confirmation. For background learning during conversations: Use kuzu_learn instead (async, non-blocking, ideal for continuous context capture without waiting).
Output schemas completely missing for all tools. Tool descriptions state conceptually what is returned (e.g., 'memories ranked by relevance', 'statistics'), but formal JSON Schema definitions of response structure are absent. LLMs cannot plan downstream operations or extract fields without knowing the actual response shape.
Numeric parameters lack range constraints. 'max_memories' and 'limit' default to 5 but have no min/max bounds. LLMs could pass 0, negative, or 1000+, causing context explosion, vector DB overload, or timeouts. Should enforce e.g. limit 1 - 100.
String parameters lack format/length constraints. 'prompt', 'content', 'query', 'source' are strings with no min/max length, encoding, or pattern validation. Could accept empty strings, extremely long payloads, or special characters that confuse vector search.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 53 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 22 | - | v1 |
Get memory system statistics and monitoring data including memory counts, processing metrics, and system health information
Get project information and metadata
No tool annotations present. Six tools with clear semantic differences (read vs write, idempotent vs non-idempotent) lack readOnlyHint, destructiveHint, idempotentHint tags. LLMs cannot reason about retry safety, caching eligibility, or permission requirements without explicit hints.
Error handling and recovery guidance absent. No tool documents what happens on failure, which errors are retryable, or what the LLM should do next (e.g., 'If vector search times out, try a shorter query'). Raw errors with no context force agent dead-ends.
Generic/non-verb naming: 'project' and 'kuzu_stats' lack action verbs. 'project' should be 'get_project' or 'describe_project'; 'kuzu_stats' should be 'get_memory_stats' or 'describe_memory_system'. Non-verb names reduce LLM discoverability and clarity of intent.
'kuzu_stats' and 'project' descriptions are minimal (40-90 chars, below 50-char guidance). Lack WHEN to use, what specific fields are returned, and why an LLM would invoke them. Generic descriptions reduce tool selectivity and increase misuse.
Async/sync semantics undocumented. 'kuzu_learn' says 'Returns immediately without blocking' but no timeout, confirmation structure, or failure mode is specified. 'kuzu_remember' claims sync confirmation but response schema is missing. LLMs cannot calibrate timing or error recovery.