Deep Code Reasoning MCP Server - Specialized analysis tool for complex codebases
The server defines 10 tools with schemas visible in src/index.ts using Zod validation. However, critical gaps severely impact quality: (1) Tool descriptions lack strategic context, they state WHAT but not WHEN or WHY to use each tool vs. alternatives. Most descriptions are 40 - 80 chars, below the 10 - 1024 baseline and well short of the 194-char average for A+ tools. (2) Parameter descriptions are sparse or generic. For example, 'entry_point' in trace_execution_path has no description explaining what 'line' and 'function_name' do. (3) Output schemas are not documented, LLMs cannot predict what fields to extract or chain to downstream tools. (4) No error recovery guidance. Tools return structured errors (ApiError, SessionError types exist in src/errors/index.ts) but the MCP handlers do not map these to actionable recovery messages. (5) All 10 tools operate on complex, poorly-described nested objects (claude_context, code_scope) with incomplete field documentation. The schema properties exist but lack descriptions for nested fields. For example, 'attempted_approaches' is documented as 'What Claude Code already tried', but 'partial_findings' has no description, the LLM must guess whether this is a list of strings, objects, or raw findings. (6) No enum constraints on repeated choice fields (e.g., analysis_type is an enum, good, but impact_types array in cross_system_impact lacks enum per-item validation in the visible schema). (7) Three tools (start_conversation, continue_conversation, finalize_conversation) manage stateful sessions but lack descriptions of session lifecycle, expiration, or retry semantics. POSITIVE: Schemas ARE present and properly typed with Zod. Naming follows verb_noun convention. Error classes suggest thoughtful error design. But execution gaps and missing descriptions drop the score significantly.
Continue an ongoing conversational analysis session
Use Gemini to analyze changes across service boundaries
Hand off complex analysis to Gemini when Claude Code hits reasoning limits. Gemini will perform deep semantic analysis beyond syntactic patterns.
End a conversation session and get final analysis results
Get the current status of an ongoing conversation session
Use Gemini to test specific theories about code behavior
Parameter descriptions missing or too generic. Nested objects (claude_context.code_scope.entry_points, analysis_type enums) have no per-field documentation. LLMs cannot infer what 'entry_points' or 'partial_findings' contain.
Tool descriptions are 25 - 40 characters, well below the 194-char average for A+ tools. Descriptions state WHAT (e.g., 'Use Gemini to perform deep execution analysis') but not WHEN (vs. other analysis tools) or HOW to interpret results.
Output schemas not documented. Tools return structured responses (session IDs, analysis results) but the MCP tool definitions do not declare return types. LLMs cannot predict chaining opportunities or extract the right fields.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 50 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 44 | - | v1 |
Use Gemini for deep performance analysis with execution modeling
Run a tournament of competing hypotheses to identify the most likely root cause
Start a conversational analysis session between Claude and Gemini
Use Gemini to perform deep execution analysis with semantic understanding
Stateful session management (start_conversation → continue_conversation → finalize_conversation) lacks lifecycle documentation. No description of session expiration, timeout behavior, concurrent limits, or retry semantics. LLMs cannot plan multi-turn sessions safely.
Error responses (ApiError, SessionError, RateLimitError) are defined in src/errors/index.ts but tool handlers do not document error cases or recovery guidance in the MCP tool definitions. LLMs will not know what to do on failure.
No pagination or result limits declared. tools like run_hypothesis_tournament may produce large tournament configs and results; no indication of max_hypotheses, max_rounds, parallel_sessions bounds in the description or response schema.
Complex claude_context object is reused across multiple tools with identical nested structure but no single schema definition. Fields like 'attempted_approaches', 'partial_findings', 'code_scope' appear in 5+ tools without clear reusable documentation.