Multi-tool MCP server providing code analysis, code graph indexing, memory graph functionality, and test running capabilities
This server has significant quality gaps. Of 23 tools, only a subset have visible input schemas in the provided code. Tool descriptions vary widely in quality but most are reasonable length (50-150 chars). However, parameter descriptions are sparse, output schemas are not documented, and error handling guidance is minimal. The code analysis tools (tools 1-7) are partially visible but lack complete schema documentation. The code graph and documentation tools (8-23) have descriptions but their actual implementations and schemas are not shown in the source provided. Per the HARD SCORING RULE: if tool definitions are inferred rather than directly visible, per-tool scores cap at 50. Most tools here fall into that category. The only tools with visible implementations are those in code_analysis/tools.py, and even there, the actual schema registration is incomplete in what was provided. The codebase shows good security practices (path validation, safe file reading, timeout handling) but fails to document output schemas, parameter constraints, or recovery guidance.
Calculate comprehensive code metrics. Returns metrics like LOC, complexity, maintainability index, etc.
Comprehensive code quality analysis using multiple tools.
Analyze code structure using AST parsing. Returns detailed information about classes, functions, imports, etc.
Analyze documentation coverage for a project using the CoverageAnalyzer.
Analyze the structure of a documentation file.
Analyze code for security vulnerabilities using bandit and custom checks.
Build the code index. If indexing is already in progress, logs a message and returns.
Output schemas not documented for any tool. LLMs cannot infer what fields to expect in responses, forcing them to guess at downstream data extraction and chaining.
Tool implementations in mcp-servers/code_graph_server/tools.py and mcp-servers/doc_validation_server/tools.py are not visible in the provided source. Tools 8-23 (13 tools, over half the suite) have descriptions but no verifiable implementation or schema. Per HARD SCORING RULE, inferred tools cap at 50.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 48 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 39 | - | v1 |
Check links in documentation files using the LinkChecker.
Find all references to a symbol in the indexed codebase.
Identify and optionally auto-fix common code issues.
Format code using specified formatter (black, autopep8).
Generate API documentation from Python code.
Generate missing docstrings for Python functions and classes.
Generate a comprehensive README template.
Generate call graph for a function.
Generate class hierarchy graph for a class.
Get the definition of a symbol in the indexed codebase.
Optimize and sort imports using isort.
Perform semantic search across the indexed codebase using embeddings.
Suggest improvements for documentation files.
Validate Python docstrings using the DocstringValidator.
Validate a Markdown file using the MarkdownValidator.
Validate a reStructuredText file using the RSTValidator.
Parameter descriptions are sparse. The visible tools (analyze_code_quality, format_code, etc.) have input schemas but descriptions for parameters like 'tools' array, 'formatter' choice, and 'direction' lack detail on valid values, constraints, or when to use each option. LLMs cannot distinguish valid inputs from hallucinated ones.
Error handling guidance is absent. No tool description explains what to do if a file is not found, analysis fails, or an external tool times out. Code shows timeout handling (30s) and safe path checks, but these safeguards are invisible to the LLM in the tool definitions.
No pagination support documented. Tools that return lists (find_references, semantic_search, analyze_doc_coverage) lack limit/offset parameters or page cursors. Responses could be large, risking context window overflow and no guidance to the LLM on handling truncation.
State mutation tools lack idempotency documentation or confirmation step. Tools like format_code, fix_code_issues, generate_docstrings, and optimize_imports all write to disk. No descriptions mention whether repeated calls are safe or if partial failures can occur.
Enum constraints missing for parameter values. 'tools' array in analyze_code_quality says it accepts [flake8, pylint, mypy, bandit] but the description does not define this as an enum. 'formatter' in format_code and 'style' in validate_docstrings/generate_docstrings also lack formal constraints.
Composition and chaining support not documented. If get_symbol_definition or find_references returns symbols, what fields (symbol_id, file_path, line_number) are included? Are those IDs valid inputs to other tools? Unclear chains force discovery detours.