Platform-agnostic AI-powered code review server via webhooks and MCP tools, supporting GitHub, GitLab, Bitbucket, and Azure DevOps
The server has 4 tools with complete input schemas and clear descriptions. However, descriptions lack depth on recovery paths and use cases. Output schemas are completely undocumented, the LLM has no visibility into what these tools return. Parameter descriptions are adequate but sparse (average ~30 chars vs baseline 72 chars). The server lacks error handling guidance, has no tool annotations (readOnlyHint), and provides minimal context on when/why to use each tool vs alternatives. This places it in fair-to-good territory, definitions are present and structured, but lack the production-grade optimization needed for A-grade scoring.
Analyze git diff and provide statistics
Review a code snippet or diff and provide AI-powered feedback with security, compilation, performance and best-practice analysis
Perform a thorough review of a complete source file — ideal for reviewing .cs, .py, .ts, .js, .go files from an IDE like Rider or VS Code
Deep security scan using OWASP Top 10 framework — finds SQL injection, XSS, hardcoded secrets, insecure deserialization and more
Output schemas completely undocumented. No visible definition of what review_code, review_file, analyze_diff, or security_scan return. LLMs cannot plan downstream calls or extract required fields without knowing response structure.
analyze_diff description is extremely sparse (22 chars). 'Analyze git diff and provide statistics' does not explain WHEN to use it vs review_code/review_file, what statistics it returns, or how the output feeds into downstream tools. Users cannot distinguish this tool's purpose.
No tool annotations (readOnlyHint). All 4 tools are marked READ_ONLY in the spec, but there is no readOnlyHint flag in the MCP tool definitions. Tools SHOULD include risk/hint metadata to guide agents on whether calls are safe to retry or have side effects.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 56 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 43 | - | v1 |
Parameter descriptions are minimal. E.g., 'focus' arrays list options (compilation, security, performance, bugs, best_practices, code_quality) but do not explain what each focus area entails or how the tool behaves when multiple focus areas conflict. baseline is 72 chars; most focus here are 20-40 chars.
No recovery guidance in error cases. Descriptions do not hint at failure modes (e.g., invalid language, missing API key, timeout). If review_code fails with 'Unsupported language', the agent has no instruction on recovery steps.
Unclear composition between tools. review_code and review_file both take 'code' + optional 'file_path' and 'provider' overrides, are they intended to be used together, or is there a clear separation of concerns? Why not combine into a single tool with a 'is_full_file' flag or 'snippet_length' detection?
Provider/model override parameters ('provider', 'model') expose implementation detail and risk token leakage if an LLM is tricked into specifying a model name containing a secret or hallucinating a non-existent provider. Consider server-side provider selection instead.