Long-term memory for Claude Code — local-first, graph-augmented, self-benchmarking
Causantic provides 4 tools with generally good descriptions (164-270 chars average) that explain WHEN to use each tool and how they differ. All tools have input parameter schemas with type declarations. However, output schemas are completely undocumented, the response format for each tool is not specified anywhere in the code. Parameter descriptions are present but lack detail on constraints, formats, and interdependencies. Tool naming is clear (verb_noun style: search, recall, predict, list-projects), but there's no evidence of input validation guidance or error recovery patterns in the tool definitions.
List all projects in memory with chunk counts and date ranges. Use to discover available project names for filtering search/recall/predict.
Predict what context or topics might be relevant based on current discussion. Walks forward through causal chains to surface likely next steps. Use this proactively to surface potentially useful past context.
Recall episodic memory — walk backward through causal chains to reconstruct narrative context. Also searches session summaries for supplementary context. Use for "how did we solve the auth bug?" or "what led to this decision?" Returns ordered narrative (problem → solution). For recent/latest session queries, use reconstruct instead.
Search memory to discover relevant past context. Uses hybrid (BM25 + vector) retrieval with entity boosting. Returns ranked results by relevance. Use this for broad discovery — "what do I know about X?" For recent/latest session queries, use reconstruct instead.
No output schemas documented for any tool. Tool descriptions mention what data is returned (e.g., 'ranked results', 'ordered narrative', 'chunk counts and date ranges'), but the actual JSON response structure is not specified. This forces LLMs to infer response fields, risking incorrect data extraction and follow-up call failures.
Parameter constraints are missing from descriptions. 'max_tokens' has no min/max bounds; 'query' and 'context' have no length guidance; 'project' and 'agent' descriptions don't clarify whether they accept names, IDs, or patterns. Absent constraints invite LLMs to pass invalid values.
Parameter descriptions are 10-30 characters on average, below the 72-char baseline. 'context' (9 chars), 'project' (6 chars), 'agent' (7 chars) are too brief for LLMs to infer correct usage. No guidance on optional vs required, no hints on how parameters interact (e.g., does 'agent' filter require 'project' first?).
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 54 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 23 | - | v1 |
No error handling patterns documented. Tools mention memory operations and queries, but there's no guidance on what happens if a project doesn't exist, if a query times out, or if memory is corrupted. Error responses should tell LLMs what to try next (retry, try a different tool, ask user), currently absent.
'reconstruct' tool is referenced in 'search' and 'recall' descriptions ('For recent/latest session queries, use reconstruct instead') but is not defined as a tool. This creates ambiguity, is 'reconstruct' a future tool, a configuration option, or an error in the description?
No idempotency guarantees or retry guidance documented. If an agent retries a 'search' or 'recall' after a timeout, will duplicate memory writes occur? Are queries idempotent? Agents need this information to plan safely.