Persistent memory and session intelligence for Claude Code: auto-tracks mistakes, decisions, and context via hooks, and mines session history for patterns and recall.
Claude Engram presents a heavily consolidated tool architecture (16 tools total from originally ~66) with significant definition quality issues. Most tools lack actionable descriptions for LLMs, parameter descriptions are often truncated or missing, and output schemas are not documented. The mega-tools (memory, work, scope, context, convention, session_mine) combine dozens of operations under single tool names with sparse descriptions of sub-operations, violating single-responsibility principle. Tools 7-9 and 16 have truncated descriptions ('description truncated in source'), making them unusable for LLM selection. No error recovery guidance, no documented output schemas, and no per-parameter validation rules. The trade-off between token efficiency (combining 66→16 tools) and definition clarity has severely compromised agent usability.
Audit multiple files for code quality issues, patterns, or security concerns.
Check Claude Engram health. Returns: status, model, memory stats.
Context operations (description truncated in source)
Convention operations (description truncated in source)
Map dependencies and imports for a file or symbol.
Generate a summary of a file's contents and purpose.
Find similar issues or patterns in the codebase.
Four tools (scope, context, convention, session_mine) have truncated descriptions marked '(description truncated in source)' and incomplete schemas with no operation enum values. These tools are completely unusable by an LLM.
Memory tool consolidates 21 sub-operations (remember, recall, forget, search, cleanup, consolidate, etc.) into a single tool with operation enum. Violates single-responsibility principle. LLM must choose among 21 operations instead of calling specific tools. This wastes reasoning tokens and increases error likelihood.
Multiple tools lack documented output schemas: memory, work, scope, context, convention, session_mine, scout_search, file_summarize, deps_map, impact_analyze, audit_batch, find_similar_issues, session_mine. Without output schemas, LLMs cannot plan downstream tool calls or extract structured results. This forces parsing unstructured text.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | F | 47 | 2026-07-28+ | v2 |
Analyze the blast radius and impact of proposed changes to a file.
Memory operations. Operations: - remember: Store a note (just content - category/relevance optional) - recall: Get all memories for project - forget: Clear project memories - search: Find by file/tags/query (file_path, tags, query, limit) - cleanup: Remove near-duplicates (same memory twice), decay, archive (dry_run, min_relevance, max_age_days) - consolidate: Merge RELATED memories in a tag group into one LLM digest, keeping the 5 most relevant and ARCHIVING the rest (tag, dry_run). Different from cleanup: that drops copies, this compresses a topic. Needs 10+ in a group; rules and mistakes are never touched; archived members stay searchable and restorable. Scope it with `tag` — a very large group (hundreds) compresses into one paragraph and loses the specifics that made each memory worth recalling - clusters: List memory clusters, or expand one (cluster_id) - add_rule: Add permanent rule (content, reason) - never decays - list_rules: Get all rules for project - modify: Edit memory (memory_id, content, relevance, category) - delete: Remove single memory (memory_id) - batch_delete: Bulk delete by IDs (memory_ids) or by category. Rules/mistakes protected from category delete. - promote: Promote memory to rule (memory_id, reason) - recent: Get recent memories newest first (category, limit) - archive: Move old inactive memories to cold storage (dry_run to preview) - restore: Bring archived memory back to active (memory_id) - archive_search: Search archived memories (query, tags, limit) - archive_status: Show hot vs archived memory counts - hybrid_search: Semantic + keyword + scored search (query, file_path, tags, limit). Best retrieval. - embed_all: Generate embeddings for all memories (enables hybrid_search) - list_mistakes: View tracked mistakes with IDs, file associations, and age - acknowledge_mistake: Archive a learned mistake so it stops appearing in pre-edit warnings (memory_id)
RARELY NEEDED — the PreToolUse hook auto-runs this before every edit. Call manually only for an explicit impact check (past mistakes, loop risk, scope violations).
Scope guard for multi-file tasks. Operations: (description truncated in source)
Search for code, files, or patterns in a directory.
RARELY NEEDED — Stop/SessionEnd hooks handle teardown automatically. Just a summary recap; memories save without it.
Mine session data for patterns, insights, and historical analysis.
RARELY NEEDED — the SessionStart hook auto-loads context every session. Call only for an explicit deep re-load (full memories + checkpoints + decisions + health).
Work tracking. Operations: - log_mistake: Record error (description, file_path, how_to_avoid) - log_decision: Record choice (decision, reason, alternatives)
audit_batch and find_similar_issues have parameters with missing type definitions (file_extensions, exclude_paths, min_severity lack 'type' attribute in schema). Parameters without types violate JSON Schema constraints and force LLMs to guess formats.
Tools like session_start and session_end emphasize 'RARELY NEEDED' but don't explain when they ARE needed or what side effects occur. No error recovery guidance if a user accidentally calls them. Descriptions lack actionable context for LLM selection.
No pagination support documented or visible in scout_search, audit_batch, or find_similar_issues. Tools that return lists should accept limit/offset and return total count to prevent context explosion.
No error handling or recovery guidance documented across any tool. No error classification (retryable, user-fixable, fatal), no examples of error responses, no guidance on what LLM should do next after failure.
Tools lack parameter validation rules. E.g., memory tool's relevance parameter accepts 1-10 but no description states this range. audit_batch shows no enum or constraint on min_severity. Without documented constraints, LLMs pass invalid values.