A multi-language MCP server with code analysis and benchmarking capabilities
Scoring was not performed
No documented output schemas. Tools return results but the structure of responses (fields, types, pagination) is not visible in the provided source code. LLMs cannot plan downstream tool usage without knowing what data comes back.
Repetitive tool design. Four analysis tools (analyze, security_scan, quality_check, secrets_scan) share nearly identical structure and parameters. This violates single-responsibility: they should be one 'scan' tool with a 'scan_type' enum parameter, or the distinctions should be clearer in naming and behavior.
Missing error handling guidance. No description of what errors these tools can return, what they mean, or how an agent should recover. E.g., what happens if 'path' is invalid? If a file cannot be parsed? Agents need recovery hints.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 39 | 2026-07-28+ | v2 |
| 2026-03-09 | C | 62 | - | v1 |
'remediate' tool accepts 'model' parameter defaulting to 'gpt-4.1' (non-existent model name). This is a hallucination risk, the LLM may assume this is valid and pass it. Should be an enum of supported models or removed entirely if server-side config is preferred.
Parameter descriptions lack actionable constraints. 'confidence' is described as 'Confidence threshold for findings (0-100)' but missing guidance: what does confidence mean? Are higher values more strict? Does 60 default mean 40% of findings are noise? LLMs need concrete semantics.
The 'remediate' tool's 'severity' parameter is optional and undefined. What severity levels are valid? (critical, high, medium, low?) This should be an enum or removed.