Memory MCP provides 13 tools with complete JSON Schema definitions and descriptions in Chinese. However, the server has critical gaps in error handling guidance, output schema documentation, and parameter validation clarity that would prevent confident production use. Tools are well-named with clear verbs (memory_cache_*, memory_get_*, memory_search_*) following verb_noun convention. Descriptions are present but often lack context on WHEN to use the tool vs alternatives, error recovery paths, and output structure. Parameter schemas are well-formed with types and enums where appropriate, but descriptions could be more explicit about constraints and formats. The STDIO-only transport and lack of tool annotations (readOnlyHint, destructiveHint, idempotentHint) limit protocol readiness. No documented output schemas for any tool, LLMs cannot plan downstream tool chains or extract structured data reliably.
NO OUTPUT SCHEMAS DOCUMENTED for any of the 13 tools. LLMs cannot plan downstream tool chains, extract specific fields, or know what data structure to expect. This violates the foundational pattern:tool requirement and pattern:response-shaper.
ERROR HANDLING AND RECOVERY GUIDANCE is entirely absent. No tool description explains what to do if a lookup fails, if an ID is invalid, if an operation conflicts, or how to retry. LLMs receive no actionable error classification (retryable vs user-fixable vs fatal). Violates pattern:recovery-guide and pattern:error-classification.
ADD OUTPUT SCHEMAS to all 13 tools. Define the response structure (fields, types, ranges) using JSON Schema. Example for memory_get_current_episode: {"type": "object", "properties": {"episode_id": {"type": "string"}, "title": {"type": "string"}, "tags": {"type": "array", "items": {"type": "string"}}, "created_at": {"type": "string", "format": "date-time"}, "message_count": {"type": "integer"}, "entity_count": {"type": "integer"}}}. Include this in tool descriptions under 'Returns:' section.
ADD ERROR GUIDANCE to every description. For each tool, explicitly state: (1) when it fails (invalid ID, no active episode, conflict); (2) what error the LLM should expect; (3) what to do next. Example: 'If candidate_id not found, returns error "Candidate not found". Verify ID via memory_get_pending() and retry.'
DOCUMENT ENTITY TYPE SEMANTICS. Add a preamble or wiki-style explanation of when to use each entity type (Decision, Preference, Concept, Habit, File, Architecture). Include in memory_add_entity description: 'Decision: significant choices made (e.g., "Use async/await for I/O"). Preference: user/system preferences (e.g., "Use TypeScript"). Concept: domain knowledge (e.g., "A circuit breaker prevents cascading failures"). Habit: recurring patterns (e.g., "Always run tests before commit"). File: source code references (e.g., "/src/auth.ts"). Architecture: system design (e.g., "Microservices with event bus").'
ADD PAGINATION METADATA to tools that return lists. For memory_get_pending, memory_recall, memory_search_by_type: document response as {"results": [...], "total_count": N, "returned": N, "next_cursor": "..." or null}. Specify: max top_k (e.g., 100), default top_k (e.g., 5), and whether results are ranked by relevance or recency.
PARAMETER DESCRIPTIONS LACK ACTIONABLE CONSTRAINTS. For example, 'entity_type' enum is listed but no description explains the semantic difference between Decision, Preference, Concept, Habit, File, and Architecture, when should an LLM pick each? Descriptions like 'tag list' (tags in memory_start_episode) do not specify format, count limits, or character restrictions. Violates pattern:constrained-input.
PAGINATION AND RESULT LIMITS are not documented. Tools like memory_get_pending, memory_recall, memory_search_by_type accept top_k but do not specify: absolute maximum, whether results are sorted, what happens if top_k exceeds available items, or whether there is a next_cursor for chaining. Violates pattern:paginated-result.
CHINESE DESCRIPTIONS limit LLM understanding in English-first contexts. While localization is valid, production MCP servers should provide English descriptions for interoperability. An English-only LLM client will not reliably understand 'entity_type: "Decision", 记忆的决策记录'.
IDEMPOTENCY NOT DOCUMENTED. Write tools like memory_confirm_entity and memory_add_entity do not state whether they are idempotent. If an LLM retries after a transient failure, will it create duplicates or safely upsert? Violates pattern:idempotent-operation.
CONFUSING TOOL OVERLAP: memory_cache_message, memory_add_entity, and memory_confirm_entity all appear to add content to memory but with unclear distinctions. Description does not explain: when should an LLM call memory_cache_message vs memory_add_entity? What is the relationship between caching a message and confirming a candidate? This violates pattern:tool-chain, tool A's output must contain the IDs tool B needs, and the roles must be obvious.
ADD TOOL ANNOTATIONS (in code and schema). Mark read-only tools with readOnlyHint: true (memory_get_*, memory_recall, memory_search_by_type, memory_stats). Mark destructive tools with destructiveHint: true (memory_close_episode, memory_deprecate_entity, memory_reject_candidate). Mark idempotent writes with idempotentHint: true if applicable. This enables LLMs to reason about side effects.
CLARIFY WRITE-OPERATION IDEMPOTENCE. State explicitly for each write tool: 'This tool is [idempotent|not idempotent]. If you retry with the same parameters, you will [get the same result|create a duplicate|throw an error].' Example for memory_add_entity: 'Not idempotent: calling twice with same content will create two separate entities. Use memory_confirm_entity to merge candidates if needed.'
SEPARATE OVERLAPPING TOOLS or document their relationship. Clarify in memory_cache_message description: 'Use this to add raw conversation messages to the current episode. The system will auto-detect entity candidates. Then use memory_confirm_entity to approve them, or memory_reject_candidate to discard.' Add similar guidance to memory_add_entity: 'Use this to manually create a named entity without parsing a message. For auto-detected candidates, use memory_confirm_entity.'
PROVIDE ENGLISH DESCRIPTIONS ALONGSIDE CHINESE. Rewrite all descriptions in English to ensure compatibility with English-first LLM clients. Keep Chinese as secondary if localization is desired. Example: 'Cache a message to the memory system (auto-detects entity candidates)' instead of '缓存一条消息到记忆系统(自动检测实体候选)'.
DOCUMENT EPISODE LIFECYCLE. Clarify in memory_start_episode, memory_close_episode, memory_get_current_episode descriptions: What is the state machine? (active → closed → archived?) Can you reopen a closed episode? Can multiple episodes be active simultaneously? Must you close before starting a new one? This prevents LLM confusion about episode boundaries.
ADD VALIDATION EXAMPLES AND CONSTRAINTS. For memory_start_episode 'tags' parameter, state: 'Array of 0-10 alphanumeric tags, each 1-32 characters. Example: ["auth", "login", "security"].' For memory_add_entity 'reason' parameter, state: 'Optional text up to 500 characters explaining why this entity was added.' Formal constraints prevent hallucinated values.
DOCUMENT 'query' SEMANTICS in memory_recall and memory_search_by_type. Are queries free-text semantic searches, regex, SQL-like? Do they search entity content, tags, metadata? State: 'Free-text semantic search using embedding similarity. Returns entities whose content is semantically nearest to the query (top_k results by relevance score).'
ADD CONFIRMATION GATE for destructive operations. For memory_deprecate_entity and memory_close_episode, consider supporting a 'dry_run' or 'confirm' parameter. Example: memory_deprecate_entity(..., dry_run=true) returns what would be deprecated without committing. This follows pattern:confirmation-request and prevents accidental data loss.