MCP server for AI-powered code analysis using Codex and Gemini
This server has well-structured tool schemas with Zod validation and comprehensive parameter definitions. However, tool descriptions are generic and lack LLM-optimization guidance. The three tools (analyze_code_with_codex, analyze_code_with_gemini, analyze_code_combined) are nearly identical in structure, violating the principle that tool names should disambiguate function. Parameter descriptions are present but lack actionable constraints for LLM reasoning. Output schemas are defined in Zod (AnalysisFindingSchema, AnalysisSummarySchema) but not surfaced in the tool registration, LLMs cannot see what fields to expect from a call. Error handling is present (Zod validation) but lacks recovery guidance.
Analyze code using both Codex and Gemini services in parallel or sequentially, with aggregated findings
Analyze code using Codex AI service
Analyze code using Gemini AI service
Tool descriptions are minimal (45 chars) and lack LLM-optimization. 'Analyze code using Codex AI service' does not explain WHEN to use this tool vs Gemini, WHAT findings it returns, or any prerequisites. Descriptions should be 50-200 chars with actionable context.
Output schemas (AnalysisFindingSchema, AnalysisSummarySchema, AnalysisMetadataSchema) are defined in src/schemas/tools.ts but NOT visible in tool registration. The MCP protocol requires tools to declare their output structure. LLMs cannot plan downstream calls or extract fields if the response schema is not documented in the tool definition.
Tool names contain no disambiguating verbs and are nearly identical. 'analyze_code_with_codex' vs 'analyze_code_with_gemini' are distinguished only by the API name. An LLM cannot infer the difference from naming alone, both start with 'analyze_code'. The 'combined' variant adds ambiguity: does it mean sequential, parallel, or aggregated? Names should clarify intent, not relegate distinction to description text.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 45 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 34 | - | v1 |
Parameter 'severity' enum is ['all', 'high', 'medium'] but description does not explain the semantics: does 'high' mean 'only high + critical' or 'high and above'? The Zod schema includes clarifying text in errorMap, but this is inaccessible to LLMs, the description must be self-contained.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) are present. All three tools are marked Risk: READ_ONLY in the metadata, but the MCP spec expects annotations in the tool definition itself. This is a CURRENT pattern (2026-07-28) that should be adopted.
Error handling lacks recovery guidance. Zod validation will reject invalid inputs (e.g. timeout < 0, severity not in enum), but the tool does not document what the LLM should do on failure: 'Invalid timeout: must be >= 0. Try with timeout=5000 or omit for unlimited.'