AI Code Surgeon MCP (OPTIMIZED): 4.9x faster, async I/O, LRU caching, Zod validation, structured logging
Server has 4 tools with explicit schema definitions visible in src/index.ts. Tool naming follows verb_noun pattern (analyze_context, think_aloud, plan_and_verify, execute_and_review), which is good. However, descriptions are present but relatively generic, they explain WHAT but lack WHEN to use and consequences. Schemas are well-structured with Zod validation, but descriptions of several parameters are vague or task-specific ('Invalid session ID format' is not helpful context for an LLM). The most critical gap: descriptions are good length (50-150 chars) but lack actionable guidance for LLM decision-making. Error handling strategy is not visible in the provided code, no recovery hints. All 4 tools have schemas, but composition is questionable: 'analyze_context' + 'think_aloud' + 'plan_and_verify' + 'execute_and_review' form a stateful workflow that requires sequential calls and session management, agents must know the exact choreography or risk failures. Output schemas are not documented.
Analyze code files for style, patterns, and issues. Creates a new session with initial analysis.
Execute approved changes and review results. Apply diffs, validate code, and optionally commit to git.
Plan code changes and verify them. Submit diffs for verification, analysis, and safety checking before execution.
Record reasoning steps in an active session. Add a thought/reasoning step to guide the AI's planning.
No output schemas documented for any tool. LLMs cannot plan downstream calls or extract data without knowing return fields.
Workflow is stateful and sequential (session-based multi-turn). Descriptions do not explain the required choreography or guard against out-of-order calls. An LLM unfamiliar with this pattern may call execute_and_review before plan_and_verify, causing failures.
Tool descriptions lack actionable guidance on WHEN to call each tool vs. alternatives. 'think_aloud' vs 'plan_and_verify' distinction is unclear, does the LLM need both? Can it skip thinking?
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | D | 57 | <=2025-11-25 | v2 |
Error handling and recovery guidance is absent. No error message examples, no 'try this next' directions. An LLM receiving a validation error on diff application has no guidance on what went wrong or how to fix it.
Parameter validation error messages appear in parameter descriptions ('Thought cannot be empty', 'At least one file is required') instead of actionable guidance. These duplicate Zod schema constraints and waste description space.
'think_aloud' is not a standard verb_noun pattern. Consider 'add_reasoning', 'record_thought', or 'add_step' to make intent clearer at a glance.
execute_and_review is a WRITE tool but the description does not explicitly warn that it modifies files and/or commits to git. Agents need to know the irreversible consequences.
No enum or constraint documentation for 'phases' parameter in plan_and_verify. What are valid phases? (e.g., 'syntax', 'safety', 'style', 'performance'?). Free-form strings invite invalid input.
No idempotency or retry semantics documented. If analyze_context, plan_and_verify, or execute_and_review are called twice with the same input, will the second call succeed or fail? Agents retry on ambiguous failures, non-idempotent tools risk duplicate side effects.