Specialized analysis tool for complex codebases - Hand off complex analysis to Gemini when Claude Code hits reasoning limits
This server demonstrates fundamental structural issues that prevent production use. While 10 tools are explicitly registered with schemas, critical gaps in descriptions, parameter documentation, and error handling guidance severely limit LLM usability. The server relies on Gemini API integration but lacks proper secret injection patterns and does not surface clear recovery paths for failures. Parameter descriptions are often missing or generic, forcing LLMs to guess intent. Output schemas are completely undocumented, there is no specification of what fields the agent should expect from successful calls. Tool compositions are unclear (e.g., start_conversation vs escalate_analysis distinction is vague). Most critically, error handling does not guide the agent toward recovery: a failed call will not help the LLM understand whether to retry, ask the user, or try a different tool.
Continue an ongoing conversational analysis session
Use Gemini to analyze changes across service boundaries
Hand off complex analysis to Gemini when Claude Code hits reasoning limits. Gemini will perform deep semantic analysis beyond syntactic patterns.
Finalize a conversational analysis session and get summary
Get the status of an ongoing conversational analysis session
Use Gemini to test specific theories about code behavior
NO OUTPUT SCHEMAS DOCUMENTED: None of the 10 tools specify what fields the LLM should expect in successful responses. The agent cannot plan downstream tool calls or extract return values. This violates the core pattern that 100% of A+ tools document return types.
MISSING OR INADEQUATE PARAMETER DESCRIPTIONS: Many parameters lack descriptions or have trivial ones (under 20 chars). E.g., 'test_approach' in hypothesis_test has no description; 'session_id' in continue_conversation has no description explaining format/lifetime. LLMs cannot infer these details from names alone.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | D | 51 | <=2025-11-25 | v2 |
Use Gemini for deep performance analysis with execution modeling
Run a hypothesis tournament to systematically test competing theories about code behavior
Start a conversational analysis session between Claude and Gemini
Use Gemini to perform deep execution analysis with semantic understanding
VAGUE/OVERLAPPING TOOL PURPOSES: The distinction between escalate_analysis, start_conversation, and continue_conversation is unclear. Both escalate_analysis and start_conversation accept claude_context and analysis_type; an LLM cannot determine when to use one vs the other. This violates the composition pattern that each tool should do exactly one thing and tool names must make the distinction obvious.
NO ERROR RECOVERY GUIDANCE: Error responses must tell the LLM what to do next (pattern:recovery-guide). The source code shows ErrorClassifier and error classes (SessionError, ApiError, etc.) but no evidence that error messages include actionable recovery hints. A 'Session not found' error should suggest alternatives, not just fail.
GEMINI_API_KEY PASSED AS ENVIRONMENT VARIABLE WITH NO SECRET INJECTION PATTERN: The code initializes deepReasoner with GEMINI_API_KEY from environment, but there is no evidence of secure server-side secret injection. If the key is exposed in tool parameters or logs, it will leak into agent traces. Patterns require credential gating via server-side vault or secure storage, never exposed to the agent.
INCOMPLETE PARAMETER CONSTRAINTS: Numeric parameters like depth_level (1-5), max_depth (1-?), max_hypotheses (2-20), and profile_depth (1-5) have some min/max constraints in schema, but many lack clear minimum/maximum specs or descriptions of what values mean. E.g., 'depth_level: How deep to analyze (1=shallow, 5=very deep)' is vague, what do levels 2, 3, 4 mean?
NO PAGINATION SUPPORT: Tools like escalate_analysis and run_hypothesis_tournament could return large result sets (hypothesis tournaments with multiple rounds), but no pagination parameters (limit, offset, next_cursor) are defined. Large result sets will blow the context window and degrade LLM reasoning.
TOOL DESCRIPTIONS TOO SHORT / LACK CONTEXT: Most tool descriptions are 40-50 characters (e.g., 'Use Gemini to test specific theories about code behavior' for hypothesis_test). While not critically short, they lack guidance on WHEN to use the tool vs alternatives, prerequisites, and expected outcomes. Baselines show A+ tools average 194 chars with clear intent/context.
MISSING CHAINING IDs IN RESPONSES: The tools deal with 'session_id' for conversations and 'entry_point' objects, but there is no specification of what IDs/references the response includes to enable downstream tool calls. E.g., if escalate_analysis returns session_id, can that session_id be passed to continue_conversation? This must be explicit in the output schema.
ANALYSIS_TYPE ENUM USED INCONSISTENTLY: Both escalate_analysis and start_conversation use analysis_type with enum values [execution_trace, cross_system, performance, hypothesis_test]. However, tools like trace_execution_path and hypothesis_test are dedicated single-purpose tools. This design suggests the analysis_type enum may be overloaded, LLMs will struggle to choose between the single-purpose tools and the generic escalate_analysis/start_conversation.