An agent memory system with a causal core — facts, temporal state, and decision→outcome causal edges on one SQLite store. MCP server for recording and retrieving causal relationships and memories.
Static source inference · medium confidence · detected: Logging
Deprecated protocol patterns detected
Summary
causal-memory demonstrates solid foundational work with 16 well-named, action-oriented tools operating on a cohesive causal inference domain. All tools have descriptions and input schemas are present with types. However, several tools suffer from incomplete parameter documentation, vague descriptions, missing output schema documentation, and weak error guidance. The server excels at composition (tools chain logically) and domain clarity (causal memory use case is well-defined), but falls short of production grade on parameter rigor and output specification. Average per-tool score: 62/100.
Tools (16)
compactwrite50/100
Consolidate the memory store (sparse-write ranking, confidence decay, semantic embeddings)
Output schemas are not documented. No visible documentation of what search_causal, search_memory, trace_cause_chain, counterfactual_query, and intervention_query return. LLMs cannot plan downstream operations without knowing field names and types.
prediction_report has empty input schema and extremely vague description ('Check whether past counterfactual advice held up'). No parameters documented. Unclear when to call this tool or what it expects. Cannot score schema rigorously.
resolve_updates description ('Resolve repeated decisions by marking supersessions') lacks clarity on what 'supersessions' means and when an LLM should invoke this. Parameter 'apply' defaults to false (preview), but description does not explain this preview-by-default workflow or the implications.
Recommendations
Document output schemas for all 16 tools. For each tool, specify: return type (object/array/string), field names, field types, and which fields are guaranteed vs optional. Example for search_causal: 'Returns an array of causal_edge objects, each with: edge_id (i64), decision (string), outcome (string), relation (string), confidence (f64), mined_from (string). Results are ordered by relevance.'
Clarify prediction_report: Add required/optional parameters and explain what data is returned. Current description is too vague. Consider renaming to validate_past_forecasts or prediction_accuracy_report for clarity.
For each search and trace tool, document whether results are paginated and whether there is a next_cursor or total_count field. If not paginated, state a hard cap: 'Returns at most 20 results; if more match, use a narrower query.'
Explain the relationship between counterfactual_query and intervention_query in their descriptions. E.g., 'counterfactual_query: find past decisions made in similar contexts and their outcomes (historical comparison). intervention_query: predict the consequences of a hypothetical decision you are considering (forecasting).'
Add a recovery guide to every tool description. E.g., for search_causal: 'If no results found, try: (1) broaden the query, (2) remove the task_tag filter, (3) lower detail_level to see more results, (4) call search_memory for cross-layer search.' This prevents silent failures.
Spec posture evidence
Inferred effective spec: <=2025-11-25.
Relies on Logging (deprecated) - log to stderr or use OpenTelemetry
Score history
Overall score trend
First recorded score · v2 rubric
63/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-23
C
63
<=2025-11-25
v2
source verified
82/100
Recording a flat fact
rememberwritesource verified80/100
Auto-extract memories from conversation text — mem0-style auto-extraction. Agent feeds raw conversation text; the system's LLM automatically extracts facts, lessons, and causal edges (caused/enabled/prevented).
resolve_updateswritesource verified70/100
Resolve repeated decisions by marking supersessions
search_causalread onlysource verified87/100
Search for past decisions and outcomes matching a query
search_factsread onlysource verified83/100
Search for facts matching a query
search_memoryread onlysource verified83/100
Search across all memory layers at once
search_patternsread onlysource verified82/100
Search for mined pattern edges in the causal store
trace_causeread onlysource verified77/100
Trace the cause of a bad outcome
trace_cause_chainread onlysource verified82/100
Trace a chain of causes backward from a bad outcome
Multiple search tools (search_causal, search_facts, search_memory, search_patterns) accept 'limit' but do not specify upper bounds or defaults in descriptions. 'limit': 5 is default, but LLMs might pass 10000. Baseline: numeric parameters should have min/max documented.
invalidate_decision and invalidate_pattern expect edge_id (i64) but do not document how an LLM obtains this ID. These tools are only useful after search or trace results return edges with IDs, but no output schema is documented to show what IDs look like or how they flow from prior tools.
remember tool description says 'mem0-style auto-extraction' without explaining what mem0 is or what 'automatic extraction' means to an LLM unfamiliar with that terminology. Description should be self-contained and not rely on external knowledge.
counterfactual_query and intervention_query descriptions do not explain the difference between them. An LLM might call both when one would suffice. Descriptions should clarify: counterfactual_query compares PAST alternatives; intervention_query FORECASTS consequences of a hypothetical decision.
No visible error handling guidance in any tool description. If search_causal finds no matches, what should the LLM do? If confidence_source is 'llm_inferred' but the system runs offline, what error is returned? Tools should guide recovery.
record_decision and record_fact do not document idempotency. Are they safe to retry? Do duplicate calls with identical input produce duplicate records or update-in-place? Agents need to know whether retries are safe.
confirm_before_execute pattern not applied to destructive tools (invalidate_decision, invalidate_pattern, compact with confidence_floor). These modify or delete data but have no confirmation step or dry-run option. An errant LLM call could purge important causal edges.
invalidate_decisioninvalidate_patterncompact
Document how edge_id flows from search/trace results to invalidate_decision/invalidate_pattern. Add an example: 'First, call search_causal to find edges. Each result includes edge_id. Pass that edge_id to invalidate_decision to mark it as wrong.'
Add dry-run parameters to invalidate_decision, invalidate_pattern, and compact. E.g., compact(confidence_floor=0.4, dry_run=true) returns what would be pruned without actually pruning. Then apply=true commits the changes. This prevents accidental data loss.
Clarify idempotency for record_decision and record_fact. If the same decision+outcome+context is recorded twice, is a duplicate created or is it merged? Document: 'Idempotent: calling with identical inputs multiple times produces one canonical record.' Or: 'Not idempotent: each call creates a new edge, but invalidate_decision can be used to mark earlier versions as superseded.'
Document permission requirements if this server requires authentication. Add scope descriptions: 'record_decision requires causal:write scope; search_causal requires causal:read scope.' This enables least-privilege agent configurations.
For remember tool, explain what 'messages' input format is expected. Is it raw text? JSON with role/content? Does the LLM need to serialize the conversation, or does the server ingest it as a stream? Current description assumes knowledge of mem0 conventions.
Add concrete examples to task_tag parameter descriptions. E.g., 'task_tag: Organize decisions by category (e.g., "cache-optimization", "auth-design", "incident-response"). Same tag + same context = comparable decision branches.' Examples prevent LLMs from inventing arbitrary tags.
Document the scope parameter for record_fact (user/session/agent). Explain: 'user: fact persists across all sessions for this user. session: fact valid only for current session. agent: fact persists only for this agent instance.' Current description is cryptic.
For confidence_source in record_decision, explain when to use each option: 'temporal: you observed the decision→outcome link directly in time. rule: rule/policy stated this causality. llm_inferred: LLM inferred the link from context (lower confidence). user_feedback: user confirmed the link.' This guides LLM selection.
Add validation constraints to detail_level enum across tools. Current descriptions show 'l0 (summary), l1 (overview), l2 (full, default)' but do not state that only these three values are valid. Make it explicit: detail_level must be one of: l0, l1, l2.
For compact tool, document pruning criteria clearly. Current description mentions 'sparse-write ranking, confidence decay, semantic embeddings' but does not explain how confidence_floor interacts with these. Add: 'Edges with confidence < confidence_floor are deleted. Remaining edges are re-ranked by recency and semantic similarity. Returns summary of pruned edges (if dry_run=true) or confirmation (if dry_run=false).'