MCP server for multi-language code graph intelligence and analysis across 25+ programming languages
CodeNavigator has 9 tools with basic schemas and descriptions, but falls short of production quality on multiple dimensions. All tools have input schemas (good), but descriptions are inconsistent, some are clear (analyze_codebase: 'Builds code graph...'), others generic (project_statistics: 'Get high-level project overview'). Parameter descriptions exist but are minimal (e.g., 'Name of the symbol to find definition for' is adequate but not LLM-optimized). No output schemas are documented anywhere. No error handling guidance is visible. No tool annotations (readOnlyHint, etc.). Tool naming follows verb_noun convention (find_definition, find_references) which is correct, but tool composition has issues: find_definition and find_references are distinct, which is good; however, find_callers and find_callees operate on 'function' parameter only, should they accept 'symbol' for generality? The analyze_codebase tool has an unusual requirement ('ALWAYS run this first') that is an anti-pattern, tool ordering should be inferred from context, not mandated in descriptions. Overall, the server is competent but mediocre, definitions are functional but lack the polish, error guidance, and schema documentation required for confident agent deployment.
Builds code graph and analyzes entire codebase. ALWAYS run this first to build the code graph
Identify refactoring opportunities by analyzing code complexity
Module relationships and circular dependencies analysis
What does this function call?
Who calls this function?
Locate where symbols are defined
Find all usages of a symbol
No output schemas documented. Tools return results but LLMs cannot know what fields to expect. This forces trial-and-error usage and breaks tool chaining.
Tool descriptions lack LLM optimization. Generic descriptions like 'Get high-level project overview and health metrics' don't answer WHEN to use the tool or distinguish it from similar tools. No dependency hints.
analyze_codebase description includes 'ALWAYS run this first', an anti-pattern. Tool ordering should be inferred from context and agent reasoning, not mandated in descriptions. This reduces agent autonomy.
Inferred effective spec: 2026-07-28+.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 55 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 29 | - | v1 |
Get comprehensive tool usage guide and best practices
Get high-level project overview and health metrics
find_callers and find_callees accept only 'function' parameter, but find_definition and find_references accept 'symbol'. This naming inconsistency forces LLMs to reason about which parameter name to use for different entity types.
No error handling guidance visible. Tools do not document what happens on failure (symbol not found, graph build failed, etc.) or how to recover. Error responses cannot guide agent retry logic.
No tool annotations (readOnlyHint, idempotentHint). All tools are read-only per the Risk field, but this is not surfaced to clients via tool metadata. Clients cannot optimize caching or predict side effects.
complexity_analysis has an optional 'threshold' parameter with default 10, but no description of the range, units, or meaning. What is a 'complexity' score? Is it cyclomatic complexity? Does 10 mean 10 branches, or percentile, or something else?