mnemo-mcp has 10 tools with mixed quality. Generic descriptions ('args dict in, envelope out') appear across pilot_* tools, providing minimal LLM guidance. Schemas are present but many descriptions are extremely terse (under 50 chars). Tools like 'pilot_capture', 'pilot_recall', 'pilot_fetch', 'pilot_reflect', 'pilot_standing_refresh', 'pilot_standing_read', 'pilot_standing_invalidate' have identical boilerplate descriptions that fail to explain WHEN to use each tool or what distinguishes them. The 'memory' and 'config' tools use action-dispatch patterns that overload responsibility into single tools. Tool annotations are present (tool_annotations=true per features), but descriptions lack the 50-200 character LLM-optimized range from production baselines. Parameters have types and basic descriptions, but lack constraints (enums for 'action', min/max for numeric k). No documented output schemas. Error handling not visible in tool definitions.
Boilerplate descriptions across pilot_* tools: 'MCP tool `pilot_capture`: args dict in, envelope out.' This tells LLMs nothing about when to use the tool, what it modifies, or how it differs from similar tools. Descriptions must be 50-200 characters and explain WHAT, WHEN, and any prerequisites.
'memory' and 'config' tools use action-dispatch pattern (single tool with 'action' parameter handling add/search/list/update/delete/export/import/stats). This violates single-responsibility principle, each action should be a separate tool (create_memory, search_memory, list_memories, update_memory, delete_memory, export_memories, import_memories, memory_stats). Composite tools confuse LLMs about when each action applies.
Replace boilerplate 'args dict in, envelope out' descriptions with domain-specific, 50-200 character descriptions. Example for pilot_capture: 'Capture a new memory with optional tags and category. Use this to store information, insights, or context the agent should remember across conversations. Returns memory_id for reference.' Explain WHAT, WHEN, WHY.
Split 'memory' tool into 8 separate tools: create_memory (add), search_memory (search), list_memories (list), update_memory (update), delete_memory (delete), export_memories (export), import_memories (import), memory_stats (stats). Each should be verb_noun with single responsibility. Agents compose them as needed, this is more discoverable than action dispatch.
Split 'config' tool into 5 separate tools: get_config (status), sync_config (sync), set_config (set), warmup_config (warmup), setup_config_sync (setup_sync). Each is clearer about what it controls.
Add JSON Schema enum constraints to remaining 'action' enums. Document the exact list: e.g. 'action: enum ["add", "search", "list", "update", "delete", "export", "import", "stats"]'.
For numeric parameters 'k', add min/max: 'k: integer, minimum 1, maximum 100 (number of results to return; default 5)'. This prevents LLM hallucination of absurd values.
Document output schemas for all tools. Example for pilot_recall: 'Returns: {"type": "object", "properties": {"memories": {"type": "array", "items": {"type": "object", "properties": {"id": {"type": "string"}, "content": {"type": "string"}, "tags": {"type": "array", "items": {"type": "string"}}, "created_at": {"type": "string", "format": "date-time"}}}}, "total_count": {"type": "integer"}}'.
Spec posture evidence
Inferred effective spec: <=2025-11-25.
Relies on Logging (deprecated) - log to stderr or use OpenTelemetry
Score history
Overall score trend
↑ 51 points across a rubric change (v1 → v2)
51/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
D
51
<=2025-11-25
v2
2026-03-09
F
0
-
v1
reversiblesource verified45/100
MCP tool ``pilot_standing_invalidate``: args dict in, envelope out.
pilot_standing_readread onlysource verified48/100
MCP tool ``pilot_standing_read``: args dict in, envelope out.
pilot_standing_refreshwritesource verified50/100
MCP tool ``pilot_standing_refresh``: args dict in, envelope out.
No enum constraints on 'action' parameter in memory and config tools. LLMs can hallucinate invalid action values (e.g. 'action: delete_all'). Must declare allowed actions as JSON Schema enum.
No documented output schemas. Tools return responses but no schema tells LLMs what fields to expect, making chaining and downstream planning impossible. Every tool must document return type.
Missing parameter constraints for numeric 'k' (number of results). No min/max specified. Baseline requires explicit range (e.g. 'k: integer, 1-100'). Unbounded numbers let LLMs request millions of results.
No error handling guidance visible. If memory_id not found in pilot_fetch, what error is returned? Are errors retryable? Do they guide next steps? Error responses must classify as retryable/user-fixable/fatal and provide recovery guidance.
Naming issue: 'pilot_standing_invalidate' is vague. Does it delete? Disable? Clear cache? Use verb_noun convention: 'delete_standing_page', 'clear_standing_page', or 'invalidate_standing_page', then clarify in description what 'invalidate' means in this domain.
pilot_standing_invalidate
Add pagination guidance: tools returning lists (pilot_recall, memory search, list_memories) should accept 'limit' and 'offset'/'cursor' parameters and return 'total_count'. Document in description: 'Results are paginated; use limit=20, offset=0 for first page.'
Add error handling guidance to tool descriptions. Example for pilot_fetch: 'If memory_id not found, returns error "Memory not found. Try search or recall to find the ID first." This is a user-fixable error, provide the search_query.'
Document which tools are idempotent vs destructive. Mark 'pilot_standing_invalidate' and 'delete_memory' as irreversible. Suggest dry-run or confirmation for destructive operations: 'This operation cannot be undone. Consider calling with a confirmation parameter.'
Add natural language resolution for identifiers where applicable. E.g., 'standing_page_key: The key for the standing page (e.g., "project-status", "team-updates"). Can be any user-friendly identifier; does not need to be a system ID.'
Clarify relationships between pilot_recall (similarity search) and pilot_reflect (reflection/synthesis). Descriptions should explain: 'recall = raw search; reflect = analysis/synthesis of results. Use recall for retrieval, reflect for deeper reasoning.'