MCP server for AutoMem: AI memory storage and recall
AutoMem provides two clearly-named tools with documented schemas and input parameters. However, descriptions are terse (recall_memory: 96 chars, store_memory: 147 chars), falling short of the 194-char baseline for A+ tools. Parameter descriptions exist but lack detail on format constraints, valid ranges, and error recovery guidance. Output schemas are not documented in the provided code. The tools are well-composed (one action each) and follow verb-noun naming convention, but descriptions do not answer WHEN to use each tool or explain dependencies between them.
Recall memories via semantic query, tags, and time windows with sorting and deduplication options
Store durable memories with type, importance, confidence, and tags; verify by recalling distinctive phrase; associate when related memory exists
Terse tool descriptions lack WHEN-to-use guidance. 'Recall memories via semantic query, tags, and time windows with sorting and deduplication options' (96 chars) does not explain when semantic search is preferred over tag-based filtering, or what the LLM should expect from deduplication. Baseline for A+ tools is 194 chars (p90=392).
Parameter descriptions lack constraints and error-recovery guidance. 'time_query' is described only as 'Time window filter (e.g., 'last 90 days')' but does not specify accepted formats, valid ranges, or what happens if the format is invalid. 'sort' accepts values like 'updated_desc' but enum values are not explicitly declared.
Output schemas are not documented in source code. The provided schema shows input structure for recall_memory and store_memory, but the response format (fields, types, pagination, total count) is not visible. LLMs cannot plan downstream tool calls without knowing what fields to expect.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 58 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 56 | - | v1 |
No pagination parameters visible for recall_memory despite accepting 'limit'. Without offset/page/cursor and total_count in response, an LLM cannot iterate over large result sets. Limit defaults and caps are not documented.
'sort' parameter lacks enum declaration. Currently documented as 'Sort order (e.g., 'updated_desc')', the example invites hallucinated values. Valid sort options should be declared as an enum (updated_desc, updated_asc, created_desc, created_asc, etc.) so LLMs select from known options.
No error-handling guidance documented. When recall_memory fails (query too vague, memory_id not found, tags invalid), the LLM has no recovery hint. E.g., 'Memory not found. Try recall_memory() with a broader query or different tags.'
store_memory lacks idempotence guarantee documentation. The description mentions 'verify by recalling distinctive phrase' but does not state whether repeated calls with identical content produce one record or duplicates. Agents retry failed calls, non-idempotent tools risk duplicate memory entries.
No dependency or sequencing documented between tools. The description mentions 'associate when related memory exists' for store_memory but does not explain if the LLM should call recall_memory first to check for related memories, or if the tool handles that internally. This creates ambiguity in multi-step planning.