MCP Server for Open Code Review - AI code quality gate for Claude Desktop, Cursor, Windsurf, VS Code Copilot
The open-code-review MCP server provides 4 tools with clear names and reasonable descriptions, but suffers from incomplete input schemas, missing output schema documentation, and insufficient parameter descriptions. Tool names follow verb-noun conventions (scan_*, explain_*, heal_*), which is good. However, parameter documentation is sparse, many parameters lack descriptions or have generic/incomplete guidance. The schema visibility is partial: input schemas are shown for some tools but output schemas are not documented anywhere in the provided source. Error handling and recovery guidance are absent. The tooling is domain-specific (code review) and well-motivated, but the interface lacks the rigor expected of production-grade agent tools.
Explain a code quality issue detected by OCR. Returns detailed explanation, category context, and fix guidance for the AI agent to act on.
Load a file's source code and prepare a repair prompt for the AI agent. The agent (you) should then apply the fix based on the issue description and suggestion. Returns the file content along with the repair context.
Scan git diff between two branches for code quality issues. Ideal for PR/MR review — only analyzes changed files and lines.
Scan a directory for AI-generated code quality issues. Detects hallucinated imports, phantom packages, stale APIs, security anti-patterns, and more. Supports TypeScript, JavaScript, Python, Java, Go, and Kotlin.
Output schemas not documented. No visibility into what scan_directory, scan_diff, explain_issue, or heal_code return. LLMs cannot plan downstream tool calls or know what fields to extract without documented return types.
Parameter descriptions are incomplete or missing context. 'path' in scan_directory and scan_diff lacks guidance on absolute vs relative paths. 'level' enum (L1/L2/L3) is named cryptically, descriptions should explain 'L1 = fast structural checks, L2 = standard + local AI analysis, L3 = deep + remote AI'. The 'languages' param in scan_directory has no validation guidance (comma-separated, but what if agent passes 'typescript;python'?).
explain_issue parameter 'severity' has no enum constraint, it's a free-form string with description 'Severity: critical/high/medium/low/info'. LLMs will pass arbitrary values. Should be an enum: ['critical', 'high', 'medium', 'low', 'info'].
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 61 | 2026-07-28+ | v2 |
No error handling or recovery guidance. If a scan fails (file permissions, syntax errors, network timeout for L2/L3), how should the LLM respond? No error schema is documented. No guidance on retryable vs fatal errors. No actionable error messages.
heal_code tool has a confusing responsibility. Description says 'Load a file's source code and prepare a repair prompt for the AI agent. The agent (you) should then apply the fix...' This mixes tool behavior (load file) with instruction to the LLM (you should apply fix). Unclear whether heal_code returns the file content, applies the fix itself, or returns a repair suggestion. Output schema is not visible.
Required parameters for scan_diff are not validated. 'base' and 'head' are git branch references, but no validation is shown. If agent passes invalid branch names, does the tool fail silently or return an actionable error?
Tool composition: scan_directory and scan_diff both return issues, but the flow to explain_issue and heal_code is unclear. Does scan_directory return issue objects that can be directly passed to explain_issue? Or must the LLM reconstruct the issue parameter as a string? No example or chaining documentation provided.
Result limits and pagination not addressed. scan_directory and scan_diff could return hundreds or thousands of issues for a large codebase. No mention of how results are limited, paginated, or sorted. LLMs receive massive unstructured lists that waste context.