A code review MCP server that analyzes code against provided context and suggests improvements using OpenAI's API
Single tool 'review_code' has a basic schema and description, but lacks critical LLM-optimization patterns. The description is 95 characters (below the 194-char production baseline and insufficient to guide LLM selection vs. similar tools). Parameters lack format constraints, enums, and dependency documentation. Output schema is not documented, the tool returns a dict with 'prompt', 'suggestions', 'is_approved', but this structure is nowhere declared or validated. Error handling is generic (ValueError wrapping) with no recovery guidance. No tool annotations (readOnlyHint, destructiveHint, idempotentHint) despite the tool calling an external API and performing I/O. The tool definition is directly visible in list_tools(), so no inference penalty applies.
Reviews code against provided context and suggests improvements
Output schema not documented. Tool returns a dict with 'prompt', 'suggestions', 'is_approved' fields, but LLMs have no visibility into this structure. Makes downstream tool chaining impossible and wastes tokens on unstructured parsing.
Parameter descriptions lack format/constraint guidance. 'context' param says only 'The context or intent of what the code should do', no hint on length, required structure, or examples. LLM will guess or hallucinate.
Tool description is too short (95 chars, vs. 194-char production baseline). Does not explain WHEN to use this tool vs. a generic linter or manual review. Missing: 'Requires OpenAI API key.' (dependency), 'Best for code intent verification, not syntax checking.' (scope).
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 38 | - | v1 |
Error handling returns generic ValueError wrapping with no recovery guidance. When OpenAI API fails or returns an error, the LLM sees 'Error analyzing code: <raw message>' and has no next step. Should return 'OpenAI API error. Retry in 5 seconds, or ask user to provide inline feedback instead.'
No tool annotations. The tool calls an external API (OpenAI), which is not idempotent (temperature=0.7 produces variable results) and could timeout. Should declare readOnlyHint=false, destructiveHint=false (no state mutation), but the variable output due to temperature is a footgun for agents retrying.
Parameters are not constrained. 'code' and 'context' accept unbounded strings. No max_length, pattern validation, or format hints. An LLM could pass a 1MB code block, causing a timeout or API quota overage.
Output parsing logic is brittle. 'is_approved' is determined by substring matching ('looks good' in analysis.lower() and 'suggest' not in analysis.lower()), this is fragile to LLM response variation and produces false positives/negatives.