Standalone MCP server for persistent AI agent memory — knowledge graph, conversation history, task tracking, and self-identity. Runs on Bun.
Forkscout Memory MCP demonstrates strong definition quality with 9 well-designed tools covering a sophisticated memory domain. Tool names follow verb_noun conventions (remember, recall, relate, task, observe, context, introspect, consolidate, forget). All tools have substantial descriptions (100-400+ chars) with clear intent statements. Input schemas are fully specified with Zod using enums, type definitions, and per-parameter descriptions. Output schemas are documented inline in implementation. Error handling includes contradiction notes and status tracking. However, there are gaps: output field names are not formally documented in response structures, parameter relationships (e.g., mode-dependent required fields) could be more explicit, and some descriptions rely on punctuation rather than clear prose structure. The 'remember' and 'task' tools show especially strong design with clear state transitions and conflict resolution strategies documented.
Memory maintenance: refresh confidence scores (sources + recency), prune low-confidence facts, remove orphaned entities, and detect duplicates. Run manually or let the server auto-run every 24h. Consolidation is IDEMPOTENT.
Working memory: push observations / facts into session context (persisted across restarts), get current context, or clear it. mode="push": add event to working memory. mode="get": retrieve current context (returns last N events). mode="clear": flush context (usually on session end). Working memory is NOT shared with recall().
Permanently remove entities, facts, relations, or exchanges from memory. mode="entity": delete entire entity + its facts + relations. mode="fact": remove one fact from an entity (by substring match). mode="relation": delete relation(s) between two entities. mode="exchange": delete conversation exchange by ID. DESTRUCTIVE — use with caution. No undo.
Self-awareness: query stats, stale entities (not accessed in N days), knowledge gaps (volatile facts not verified recently), and relation graph health. Use to diagnose memory quality and plan consolidation.
Record a conversation exchange into long-term memory. Auto-infers importance from keyword signals (fix, bug, decision, learned, etc.). New facts are NOT auto-pushed to recall() — call context(mode="push") to surface them. Optional: tags= (scoped search), sessionId= (to group multiple exchanges).
Output response schemas not formally documented. Tools return structured content (text + metadata) but per-field contracts are implicit. LLMs cannot reliably extract specific fields from responses.
Mode-dependent required parameters not explicitly documented. E.g., 'recall' requires 'name' for mode=entity/history but this dependency is buried in prose. Schema should use conditional validators or explicit mode/param relationship tables.
Parameter format constraints stated informally. E.g., 'weight: 0 - 1' and 'priority: 0 - 1' are described in prose but not enforced via JSON Schema min/max. LLMs may pass invalid values outside bounds.
Inferred effective spec: 2026-07-28+.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 68 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 47 | - | v1 |
Multi-modal retrieval across all memory. mode="search" (default): BM25 search across entities + conversation history. mode="entity": get one entity by exact name (requires name=); optional query= filters facts. mode="history": belief evolution — what was corrected and when (requires name=). mode="relations": knowledge graph edges (optional name= restricts to one entity). mode="exchanges": browse or search conversation history. Call without query= to list all exchanges newest-first (use offset= + limit= to paginate). Call with query= to filter by keyword. Response always includes total count and pool size. mode="timeline": chronological activity feed — recent exchanges + entity accesses merged by time.
Create a typed knowledge-graph relationship between two entities. Optionally weight= the relation (0–1, default 0.5) and provide evidence= to strengthen it. Relations auto-inferred on remember() — use this only for explicit, intentional edges.
Store, update, or supersede facts on a named entity in the knowledge graph. If the entity exists, new facts are merged and contradictions auto-resolved. To UPDATE: pass supersede="old fact substring" + facts=["replacement"]. To REMOVE: pass facts=[] + supersede="substring to drop". Use entity name "Forkscout Agent" to record self-observations. RULE: Call recall(mode="search") first — avoid creating duplicate entities.
Executive memory: start, update, complete, or abort tasks. mode="start": create a task (or resume existing similar one). mode="update": touch or patch (priority, budget). mode="complete": mark done + auto-save success record. mode="abort": cancel + auto-save failure post-mortem. mode="list": get all active tasks (or filter by status=). mode="summary": sprint-format summary of running/paused tasks (for agent context). All terminal tasks auto-pruned (keep newest 50).
'context' tool description is terse (78 chars). Lacks detail on what 'working memory' means vs long-term memory (recall) or how it persists across restarts. Under pattern:tool-description baseline (194 chars).
No explicit recovery guidance in error scenarios. E.g., if 'supersede' substring not found on 'remember', response is generic. Should guide LLM: 'Substring not found. Call recall(mode=entity, name=...) to see available facts.'
Task completion/abort does not document what success/failure records look like. E.g., mode=complete auto-saves but response doesn't show the structure of the saved success record. Breaks chaining assumptions.