mnemo-mcp shows moderate definition quality with significant gaps. The server has 10 tools with mixed naming clarity, description coverage, and schema completeness. Tool names like 'memory' and 'pilot_*' lack action verbs, making intent ambiguous. Descriptions exist but are often generic or abbreviated (e.g., 'MCP tool pilot_recall: args dict in, envelope out' provides minimal guidance). Input schemas are visible but lack granularity, many parameters are untyped or loosely constrained. The 'memory' tool bundles 9 distinct actions (capture, recall, fetch, list, update, delete, export, import, stats) into a single tool, violating single-responsibility. No output schemas are documented. Error handling is not evident from the definitions.
Tools (10)
configwriteauthsource verified45/100
Configure mnemo server: status (show configuration), sync (trigger sync), set (update settings), warmup (pre-download models), setup_sync (authenticate cloud sync)
helpread onlysource verified63/100
Full documentation on demand covering all memory operations and configuration
Tool 'memory' bundles 9 distinct operations (capture, recall, fetch, list, update, delete, export, import, stats) into a single tool. This violates single-responsibility and forces LLMs to reason about which action to invoke, increasing errors.
Tool 'config' bundles 5 actions (status, sync, set, warmup, setup_sync) into a single tool without clear differentiation. LLMs struggle to choose the right action.
All 'pilot_*' tools (5 tools) have generic boilerplate descriptions like 'MCP tool pilot_recall: args dict in, envelope out' that provide no actionable guidance. Descriptions must explain WHAT the tool does, WHEN to use it, and key parameters.
Refactor 'memory' tool into separate single-action tools: create_memory (capture), search_memories (recall), get_memory (fetch), list_memories, update_memory, delete_memory, export_memories, import_memories, get_memory_stats. Each tool then has one clear responsibility and name (verb_noun).
Refactor 'config' tool into: get_config (status), sync_config (sync), set_config (set), warmup_models (warmup), setup_cloud_sync (setup_sync). This improves clarity and discoverability.
Replace all boilerplate descriptions like 'MCP tool pilot_recall: args dict in, envelope out' with actionable descriptions that explain: (1) what the tool does, (2) when to use it vs alternatives, (3) key parameters and constraints. Target 50 - 200 characters per Arcade baseline. Example for pilot_recall: 'Search stored memories using semantic similarity. Returns up to k=5 results ranked by relevance. Use after capture to retrieve specific knowledge.'
Rename 'pilot_*' tools to drop the 'pilot' prefix and use verb-first naming: 'capture_memory', 'search_memories', 'fetch_memory', 'reflect_on_context', 'update_standing_page', 'get_standing_page', 'invalidate_standing_page'. This makes names self-documenting and helps LLMs parse intent.
Add enum constraints to 'action' parameters in 'memory' and 'config' tools. Example: action: {type: 'string', enum: ['capture', 'recall', 'fetch', 'list', 'update', 'delete', 'export', 'import', 'stats'], description: 'Memory operation: capture=store, recall=search, fetch=retrieve by id, list=enumerate, update=modify, delete=remove, export=serialize, import=deserialize, stats=usage metrics'}.
Spec posture evidence
Inferred effective spec: <=2025-11-25.
Relies on Logging (deprecated) - log to stderr or use OpenTelemetry
Score history
Overall score trend
First recorded score · v2 rubric
52/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-23
D
52
<=2025-11-25
v2
MCP tool pilot_reflect: args dict in, envelope out
Tool names do not consistently start with action verbs. 'memory' and 'config' are nouns; 'pilot_*' have unclear prefixes. Verb-first names (capture_memory, get_config) help LLMs infer intent without reading descriptions.
No output schemas are documented for any tool. LLMs need to know what fields to expect in responses to plan downstream calls and extract data. This blocks composition and causes LLMs to guess field names.
Input schemas lack enum constraints for enumerated fields. 'action' parameters in 'memory' and 'config' tools are free-form strings, inviting hallucinated invalid actions. Enums are self-documenting and prevent invalid invocations.
No parameter constraints documented for numeric fields like 'k' (pilot_recall, pilot_reflect). Should specify valid range (e.g., 1 - 100) to prevent LLMs from passing absurd values (k=1000000).
Tool descriptions lack context on when to use one tool vs another. E.g., what is the difference between 'pilot_recall' (search) and 'pilot_reflect' (synthesis)? Without clear delineation, LLMs conflate similar tools.
No error handling guidance visible in tool definitions. Tools should document expected errors (e.g., 'memory_id not found') and recovery steps (e.g., 'Use pilot_recall to search for the memory').
Add range constraints to numeric parameters. Example for 'k' in pilot_recall: {type: 'integer', minimum: 1, maximum: 100, default: 5, description: 'Number of results to return (1 - 100)'} to prevent LLMs from requesting thousands of results.
Document output schemas for all tools. Example for pilot_recall: {type: 'object', properties: {results: {type: 'array', items: {type: 'object', properties: {memory_id: {type: 'string'}, content: {type: 'string'}, score: {type: 'number'}, timestamp: {type: 'string'}}}}, total: {type: 'integer'}}}. This enables LLMs to plan chains and extract fields.
Add error handling guidance to tool descriptions. Example for pilot_fetch: 'Retrieve a memory by ID. Returns memory object or error 404 if not found. If you have only a partial memory, use search_memories first. Errors: 404 (memory not found, try search_memories), 401 (unauthorized, check credentials)'.
Clarify what 'standing pages' are in descriptions. Example for pilot_standing_refresh: 'Materialize a standing page (persistent, up-to-date summary) by answering a question using relevant memories. Returns page key and content. Use to maintain dynamic knowledge synopses.'
Add parameter descriptions for 'key' in standing page tools explaining expected format and valid values (e.g., 'Alphanumeric page identifier (2 - 50 chars), e.g. team-status, project-roadmap').
Add guidance on idempotency, state mutations, and retry safety. Example: 'capture_memory is idempotent, calling twice with same content returns same memory_id. delete_memory is irreversible, no undo. Reflect tools are read-only and safe to retry.'
Consider documenting which tools require authentication (e.g., setup_cloud_sync) vs which work offline, to help agents plan around permission failures.