Local-first memory and assertion ledger for AI coding agents. Confidence-weighted, contradiction-aware, token-budgeted context over MCP.
Engram has 23 tools with complete input schemas and descriptions. Naming is verb-forward (remember, search, get, list, query, rebuild, find, handoff, receive, replay, diff, append). Descriptions are present and contextual (avg ~120 chars), meeting the 10-1024 char baseline. However, several tools lack output schema documentation, parameter descriptions are sometimes generic, and error handling guidance is minimal. No tool annotations (readOnlyHint/destructiveHint) despite clear risk classifications. Tools are well-composed (single responsibility) and most accept natural identifiers (project names, queries). Validation logic exists in mcp-tools.js but is not surfaced in schema constraints.
Search across ALL projects semantically. Groups results by project with relevance scores. Great for finding related work across the entire codebase.
Find duplicate or near-duplicate sessions using embedding similarity. Returns pairs of sessions above the similarity threshold.
Get pre-compiled context bundle for a specific project. Includes description, tech stack, recent sessions, and key concepts.
Get a summary of the knowledge graph.
Get full session details by project and session ID.
Get overall Engram statistics.
Get top topics/tags from Engram with session counts.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) despite clear risk classifications. Tools marked WRITE (remember, rebuild_index, ledger_ingest, handoff, session_append_event) should declare destructiveHint=true; READ_ONLY tools should declare readOnlyHint=true. This prevents agents from accidentally treating read-only tools as safe to retry or destructive tools as idempotent.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | D | 59 | 2026-07-28+ | v2 |
Prepare context for handoff to another agent.
Write an assertion to the ledger. Creates new or reinforces existing assertions.
Query assertions from the ledger by topic or pattern.
Select relevant assertions from the ledger to include in context.
Get statistics about the assertion ledger.
List all projects indexed in Engram with session counts.
Semantic search across all Engram sessions. Finds sessions by meaning, not just keywords. Use this to find relevant past work, learnings, and context.
Look up a concept in the knowledge graph. Returns related concepts and connection strengths.
Rebuild Engram indexes (bloom, git, embeddings).
Receive context from another agent.
Get the most recent sessions across all projects.
Save a memory or session to Engram. Call at end of session, after completing a feature, or when recording a decision.
Keyword search across sessions by summary and topics.
Append an event to an existing session.
Compare two sessions to see what changed.
Replay a session to see the sequence of events and decisions.
Output schemas not documented. Tools return structured data (e.g., neural_search returns {query, total_matches, decay_enabled, results}, get_bundle returns context bundles) but the response schema is not visible in the tool definition. LLMs cannot plan downstream tool calls or extract fields without knowing the output structure.
Error handling lacks recovery guidance. validateRememberInput() returns structured errors (ENGRAM_ERR_SUMMARY_REQUIRED, ENGRAM_ERR_TOPICS_REQUIRED) but tool descriptions do not explain what to do on failure. E.g., remember should state: 'If validation fails, check that summary is 1 - 1000 chars and topics is a non-empty array of 2 - 8 tags.'
Parameter descriptions are generic or missing context. E.g., 'limit' appears in 8 tools with identical description 'Maximum results to return (default: 10)' but does not explain the impact of high limits on token usage or performance. 'query' in neural_search says 'searches by semantic meaning' but does not explain when to use neural_search vs search_sessions vs cross_project_search.
No pagination guidance for large result sets. Tools like list_projects, get_topics, and cross_project_search accept 'limit' but do not document whether results are sorted, whether a next_cursor or offset is available, or what happens when limit is exceeded. This risks context window exhaustion.