Local-first memory and assertion ledger for AI coding agents. Confidence-weighted, contradiction-aware, token-budgeted context over MCP.
Engram provides 23 tools with complete JSON Schema input definitions and descriptions for all tools. Most tool names follow verb_noun convention (remember, neural_search, get_bundle, search_sessions, list_projects, query_concept, rebuild_index, ledger_ingest, ledger_query). Descriptions are present and substantive (average ~130 chars, within the baseline 194-char median). However, there are significant gaps: (1) Parameter descriptions are inconsistent, many lack actionable detail about format, constraints, or when to use them. E.g., 'query' params say 'The search query' or 'Keyword query to search for' but do not specify max length, character restrictions, or examples. (2) Output schemas are NOT documented in the source code, we can see input schemas but have no visibility into response structures, field types, or pagination patterns. This forces LLMs to guess what fields exist in responses. (3) Error handling is not evident, no examples of error messages, recovery guidance, or categorization (retryable vs fatal). (4) Several tools have trivial or generic descriptions: get_stats (13 chars), get_graph_summary (14 chars), session_append_event (25 chars). (5) Parameter naming inconsistencies: some tools use 'query', others use 'pattern' for search-like operations. (6) No indication of which operations are idempotent, destructive, or read-only in the tool descriptions themselves, only inferred from the Risk field in the metadata. (7) Missing output pagination guidance: neural_search, cross_project_search, find_duplicates return multiple results but no documented limit enforcement or next_cursor pattern.
Search across ALL projects semantically. Groups results by project with relevance scores. Great for finding related work across the entire codebase.
Find duplicate or near-duplicate sessions using embedding similarity. Returns pairs of sessions above the similarity threshold.
Get pre-compiled context bundle for a specific project. Includes description, tech stack, recent sessions, and key concepts.
Get a summary of the knowledge graph (top concepts, connection density).
Get full session details by project and session ID.
Get Engram statistics (total sessions, projects, topics, graph nodes).
Get top topics/tags from Engram with session counts.
Output schemas are not documented. Input schemas are complete with types and descriptions, but response structures are invisible. LLMs cannot plan downstream tool calls or extract fields without trial and error.
Parameter descriptions lack actionable constraints. Many search/query parameters say 'The search query' without specifying max length, allowed characters, or performance implications (e.g., does a very long query timeout?). Field names like 'query' vs 'pattern' are used inconsistently across similar tools.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 64 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 52 | - | v1 |
Create a handoff context summary for another agent or session.
Write an assertion to the ledger. Creates new or reinforces existing assertions.
Query assertions from the ledger by pattern.
Select high-confidence assertions to include in context.
Get statistics on the assertion ledger.
List all projects indexed in Engram with session counts.
Semantic search across all Engram sessions. Finds sessions by meaning, not just keywords. Use this to find relevant past work, learnings, and context.
Look up a concept in the knowledge graph. Returns related concepts and connection strengths.
Rebuild Engram indexes (bloom, git, embeddings).
Receive and ingest a handoff context from another agent.
Get the most recent sessions across all projects.
Save a memory or session to Engram. Call at end of session, after completing a feature, or when recording a decision.
Keyword search across sessions by summary and topics.
Append an event to an existing session.
Compare two sessions to see what changed.
Replay a session step-by-step to understand the decision path.
Tool descriptions are often generic or too brief (<50 chars) and do not explain WHEN to use this tool vs similar ones. E.g., 'get_stats' (13 chars), 'get_graph_summary' (14 chars). LLMs cannot disambiguate between neural_search, search_sessions, cross_project_search, and query_concept without clearer guidance on use cases.
No error handling documentation. No examples of error messages, recovery guidance, or how to handle partial failures (e.g., what if embeddings are not ready for neural_search?). Errors likely return raw messages without actionable next steps.
Result limits not enforced or documented. neural_search, cross_project_search, search_sessions, find_duplicates all return multiple results but descriptions lack guidance on limit enforcement, pagination, or when results are truncated. No mention of next_cursor or offset patterns.
Destructive operations (remember, rebuild_index, ledger_ingest, receive_handoff, session_append_event) lack confirmation or dry-run hints. No guidance that these are irreversible or how to preview changes before commit.
Tool naming inconsistencies in semantic search: neural_search, cross_project_search, search_sessions, query_concept, ledger_query all perform search-like operations but use different names ('neural_search' vs 'cross_project_search' vs 'search_sessions'). LLM reasoning cycles wasted deciding which to use.