MCP server providing persistent memory across conversations
Munin Memory is a 23-tool knowledge base MCP server with comprehensive tool coverage, strong naming conventions, and detailed parameter schemas. However, it suffers from inconsistent description quality, missing output schema documentation, and weak error handling guidance. Most tools follow verb_noun naming (write, read, query, orient, resume, extract, narrative, commitments, patterns, handoff, log, list, delete, attention, insights, audit_history, consolidate, retrieval_feedback, dashboard, maintenance, status_update, read_batch, get). Parameter descriptions are present and mostly well-formed with type constraints (enums for search_mode, entry_type, detail, operation). However: (1) Tool descriptions vary wildly in specificity (some are 20-30 chars, below the 50-100 baseline for LLM optimization); (2) Output schemas are almost entirely absent from the visible code, we can infer from function names what write() and read() return, but complex tools like query(), narrative(), extract(), insights(), and consolidate() have no documented return structure; (3) Error handling is not visible in tool definitions, no recovery guidance, no actionable error messages shown; (4) Destructive operations (delete, consolidate, maintenance) lack confirmation/dry-run parameters in the schema shown, though the description for delete mentions two-step confirmation; (5) No tool annotations (readOnlyHint, destructiveHint, idempotentHint) visible in registration.
Get items requiring attention (unreviewed, overdue, or high-priority)
Retrieve the audit log of changes to the knowledge base
List and manage commitments tracked in the knowledge base
Synthesize and consolidate entries in a namespace (uses LLM for summarization)
Get a comprehensive dashboard view of the knowledge base state
Delete entries from the knowledge base (requires two-step confirmation)
Output schemas are not documented for most tools. Tools like query(), extract(), narrative(), insights(), consolidate(), dashboard() have no visible return type specification in tool registration. LLMs cannot plan downstream composition or extract specific fields without knowing the response structure.
Tool descriptions vary significantly in quality. Many are 50-70 characters (below the 100-character baseline for LLM optimization). Examples: 'orient' (79 chars), 'resume' (83 chars), 'patterns' (76 chars), 'attention' (80 chars). Descriptions should clearly state WHAT, WHEN, and WHY the agent should use the tool.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 67 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 38 | 2025-03-26+ | v1 |
Extract structured suggestions from the knowledge base (related entries, patterns, topics)
Retrieve a specific entry by namespace and key
Create a structured handoff summary for transferring context
Get AI-powered insights and analysis about an entry or topic
List entries in a namespace with optional filtering and pagination
Add an event log entry to record actions and observations
Run maintenance operations on the knowledge base
Generate a narrative summary from log entries with timeline and sources
Get an overview of the knowledge base state, organization, and key contexts
Identify recurring patterns and themes in the knowledge base
Search the knowledge base using lexical, semantic, or hybrid search modes
Retrieve a single entry by ID
Retrieve multiple entries by ID in a single call
Get a summary of open loops, suggested reads, and actionable items
Log feedback on retrieval quality to improve search results
Update status of an existing entry without changing core content
Store state, logs, and commitments in the knowledge base
No tool annotations visible (readOnlyHint, destructiveHint, idempotentHint). Agents cannot determine which tools are safe to retry or which modify state. Destructive tools (delete, consolidate, maintenance) should be marked with destructiveHint=true and require confirmation. Read-only tools should be marked with readOnlyHint=true.
Error handling and recovery guidance are not documented in tool descriptions. Tools provide no guidance on what to do if a call fails (retryable? user-fixable? fatal?). For example, delete() has no guidance on what happens if the confirm_token is invalid, or what should be retried.
Destructive operations lack explicit confirmation/dry-run parameters in schemas. consolidate() accepts a 'preview_token' (good), but delete() should show a dry-run or preview step before actual deletion. maintenance() has a dry_run parameter (good) but it is optional and may default to false, risking accidental data loss.
Some tool names are vague or lack clear action verbs. 'orient' (to what?), 'resume' (resume what?), 'extract' (extract what?). These names do not clearly convey what happens when called. Compare to 'get_knowledge_base_overview', 'list_pending_items', 'extract_related_entries', more explicit.
Parameter descriptions sometimes lack constraint details. For example, 'limit' and 'offset' parameters do not specify min/max bounds. 'tags' arrays do not specify max length. 'content' in write() has no length limit documented. These unconstrained params let LLMs pass absurd values (limit=999999, content with megabytes of text).
No documented pagination or result limits. Tools that return lists (list, query, commitments, patterns, audit_history, attention, insights) should document max result size and whether pagination is supported. Large result sets will exhaust context windows and degrade LLM reasoning.