Long-term memory for AI agents, powered by SQLite, FTS5, and user-editable categories
ContextKeep V2.1 presents a mixed picture. The server provides 16 well-named, verb-first tools with consistent descriptions covering memory management and category operations. All tools have descriptions (10 - 150 chars, within baseline). However, critical gaps emerge in schema completeness and output documentation. Input schemas are visible and properly typed for all tools, but output schemas are COMPLETELY UNDOCUMENTED, the code returns formatted string responses without declaring the shape of structured data. Parameter descriptions are present but minimal (15 - 40 chars), lacking guidance on constraints, formats, or dependencies. Error handling is weak: functions return plain strings without actionable recovery steps or error classification. The server lacks security hardening (no audit trails, no permission gates, no rate limiting documented). Tool composition is generally sound (single responsibility, verb_noun naming), but several tools accept string-formatted lists (e.g., 'categories' as comma-separated string) instead of arrays, forcing LLMs to construct ad-hoc serialization. Based on the 549-tool baseline (average description 194 chars, param annotation 72 chars), this server's descriptions are BELOW baseline and lack LLM-optimization depth.
Create a new live category.
Delete an empty category or reassign its memories to another category.
Delete a memory permanently by key.
Export all memories as JSON, including categories and legacy tags.
Get ContextKeep version, backend, schema, tools, and migration status.
Get edit history for a memory.
Get memory and category statistics.
Output schemas completely undocumented. All 16 tools return formatted strings; LLM does not know the structure of underlying data (e.g., memory dict fields, category dict fields, history array shape). This violates pattern:tool and pattern:response-shaper, LLM cannot reliably extract or chain data without parsing unstructured text.
Multiple tools accept comma-separated string parameters (e.g., 'categories', 'category') instead of typed arrays or enums. Forces LLM to construct string serialization manually (category1,category2) and invites delimiter escaping bugs. Per pattern:constrained-input, should use enum or array types.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 69 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 47 | - | v1 |
List live memory categories with counts and descriptions.
List memories, optionally filtered by category. V2 replacement for list_all_memories.
List recently updated memories.
Merge one category into another.
Retrieve a memory by exact key.
Full-text search memories, optionally scoped to a category.
Store or update a memory with comma-separated category assignments.
Reassign categories for an existing memory.
Rename or edit category metadata by category id.
No error guidance. Tools return plain strings ('Memory not found: {key}', 'Category not found', etc.) without recovery hints, categorization, or actionable next steps. Per pattern:recovery-guide, errors must guide the agent: 'User not found. Try search_categories() first.'
Destructive operations (delete_memory, delete_category, merge_categories) lack confirmation/dry-run patterns. Per pattern:confirmation-request, irreversible actions should require approval or offer a preview step before committing. LLM can delete without user consent.
Parameter descriptions are below baseline. Average 25 chars vs. baseline 72 chars. Many lack format constraints, enums, min/max bounds, or context. E.g., 'Comma-separated category assignments' does not explain: max length per category name, max number of categories per memory, what happens if category doesn't exist, whether duplicates are allowed.
Tool descriptions average 65 chars vs. baseline 194 chars. Descriptions are present but minimal, lacking WHEN to use, prerequisites, or trade-offs. E.g., 'List memories, optionally filtered by category' does not explain why to call list_memories vs. list_recent_memories, whether results are paginated, or what happens with large datasets.
No pagination support on result-returning tools. list_memories and list_recent_memories accept a limit but do not return pagination cursors (next_token, total_count) or support offset/page parameters. export_memories returns all records in one response (context explosion). Per pattern:paginated-result, unbounded lists invite context exhaustion.
No security hardening visible. No audit trails, permission checks, rate limits, or secret injection patterns documented. Per pattern:secret-injection and pattern:audit-trail, production systems must log who called what and validate permissions. Per pattern:scope-declaration, tools should declare required scopes (e.g., read:memory, write:memory, delete:memory).
'categories' parameter in store_memory and update_categories uses string serialization (comma-separated) rather than array type. This is a design smell, if a category name contains a comma, parsing fails. Per pattern:constrained-input, should use array of strings with enum or validate against known category IDs.
No response field naming consistency for chaining. If list_memories returns 'key' and retrieve_memory expects 'key', that's good. But get_memory_stats returns 'per_category_counts' while update_category returns 'category', field names vary. Per mxe:response-field-naming, consistent naming prevents LLM confusion when chaining tools.