EverMem Plugin for Claude Code - automatic memory recall from past sessions. Exposes memory search tool for Claude to find relevant context from past sessions.
Single tool with acceptable naming and schema but significant gaps in description quality and depth. The tool name 'evermem_search' follows verb_noun convention and clearly indicates a search action. Input schema is properly structured with JSON Schema types and has basic descriptions. However, the tool description (96 chars) falls in the acceptable range but lacks specificity about what 'memories' are, when to use this vs alternatives, and what format results return. Parameter descriptions are minimal, 'query' and 'limit' are present but could be more explicit about acceptable formats, constraints, and examples of good queries. No documented output schema provided, making it unclear what fields the LLM should expect in returned summaries. The rubric baseline shows average tool descriptions are 194 chars and param annotations 72 chars; this tool is significantly below average. No error handling documentation provided, LLMs won't know how to handle 'no results found' or invalid queries. Risk annotation (READ_ONLY) is present but tool lacks error guidance, recovery hints, or composition context (e.g., how results chain to other tools).
Search past conversation memories. Returns summaries with dates and relevance scores. Use when user asks about previous work, decisions, or context from past sessions. Params: query (required), limit (default: 10, max: 20)
Missing output schema documentation. LLMs cannot determine what fields to extract from results (e.g., are summaries plain text or structured? what fields are in each result object?). No documented return type violates pattern:tool baseline (100% of A+ tools have documented return types).
Tool description is too brief (96 chars vs baseline 194 chars). Lacks WHEN to use this tool, what 'memories' means in domain context, and what format results return. Does not answer the three critical questions: What does it do? When should the LLM call it instead of alternatives? What does it return?
Parameter descriptions lack actionable detail. 'Search query - use keywords, topics, or questions' is helpful but provides no constraints on format, length, or valid input patterns. Does not specify minimum/maximum query length, what happens with empty queries, or whether boolean operators are supported. Baseline: 72 chars per param annotation, these are underdescribed.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 59 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
No error handling or recovery guidance. Tool description does not tell LLM what to do if query returns no results, if limit exceeds max (20), or if query is malformed. Pattern:recovery-guide requires error responses to say 'what to do next'; this tool has no error documentation at all.
Tool composition unclear. Does 'evermem_search' return context that chains to other tools or functions? Are returned memory IDs expected to be passed to follow-up calls? No documentation of what downstream actions are enabled by this tool's output.