Semantic code intelligence for AI-native development. Exposes MU capabilities as MCP tools for code analysis, search, graph traversal, and codebase compression.
MU MCP Server provides 14 well-intentioned tools for semantic code analysis. Tool naming follows verb_noun convention (mu_search, mu_expand, mu_read) which is appropriate for internal DSL. Descriptions are present for all tools and generally clear (median ~120 chars, within baseline range). Parameter schemas are fully specified with types (usize, u8, bool, string, arrays) and descriptions. However, several critical gaps reduce overall score: (1) Output schemas are NOT documented, the tools return markdown strings with no formal field definitions, forcing LLMs to parse unstructured text; (2) No error handling guidance, tools lack actionable recovery instructions or error classification; (3) Parameter defaults are underspecified, many accept optional params without explicit default behavior documented; (4) No pagination/limits guidance beyond inline comments; (5) Tool composition issues, mu_review combines mu_diff + mu_impact + mu_audit but the integration is not formally documented; (6) Security considerations for WRITE tools (mu_configure, mu_enrich, mu_bootstrap) lack permission gates or audit trail declarations; (7) STDIO-only transport hard-caps protocol readiness at 50, preventing remote accessibility. Despite these gaps, the server demonstrates solid foundational quality: parameter lenient deserializers show care for client tolerance, tool descriptions explain WHAT and WHEN to use each tool, and the semantic code analysis domain is well-scoped.
Comprehensive codebase audit with built-in rules and project-local custom rules. Detects violations, code smells, and missing docs.
Initialize and build MU database. Parses codebase into semantic graph, computes importance scores, enables all other tools.
Compress entire codebase into hierarchical token-efficient representation for LLM consumption. Degrades detail by importance to fit token budget.
Configure MU analysis: set service classifications, domain concepts, priority nodes, test handling. Supports discovery mode and incremental enrichment.
Semantic diff between git references. Shows changed nodes, their impact, and complexity deltas.
Store or retrieve enrichment data (summaries, analyses) for nodes. Enables model-generated annotations to improve future analyses.
No output schemas documented. Tools return formatted markdown strings with no formal response field definitions. LLMs must parse unstructured text, losing the ability to reliably extract structured data for chaining.
No error handling guidance. Tools lack actionable recovery instructions (e.g., 'If project not found, call mu_search first'). Bare error messages provide no path forward for LLM agents.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 66 | 2026-07-28+ | v2 |
Expand from seed nodes by traversing edges in the dependency graph. Returns connected nodes and edges up to specified depth.
Find symbols by exact name, qualified name, dotted suffix, or node ID. Returns matching definitions ranked by type and importance.
Analyze downstream impact of changing a symbol. Uses BFS to find all code that depends on the target.
Pack high-value code context into token budget. Optionally task-aware or node-specific. Returns grouped by file or flat.
Read detailed content of nodes. Supports multiple detail levels: signature, summary, source, or full code.
Review changes: semantic diff + impact analysis + audit findings + risk score. Combines mu_diff, mu_impact, and mu_audit.
Search nodes in the codebase by natural language query or symbol name. Returns ranked results with confidence scores.
Find suspicious code patterns: high complexity, many parameters, missing docs, code smells.
WRITE tools (mu_configure, mu_enrich, mu_bootstrap) lack permission gates and audit trail declarations. No documentation of what access control is required or how changes are logged.
mu_configure accepts JSON string in 'corrections' param but no schema/validation rules documented. LLMs cannot discover valid correction keys (service_classifications, domain_concepts, etc.) without trial-and-error.
Parameter defaults are under-specified. 'limit' in mu_search defaults to 10, 'depth' in mu_expand defaults to 1, but the descriptions do not make these defaults explicit. LLMs may assume different behavior.
mu_sus tool name is cryptic ('sus' likely means 'suspicious'). Expand to 'find_suspicious_code' or similar to make intent clear without requiring context.
mu_review is documented as combining mu_diff + mu_impact + mu_audit, but the composition contract is not formal. No documentation of which output fields come from which sub-operation, risking LLM confusion.
Tools returning lists (mu_search, mu_expand, mu_audit results) do not document pagination. No mention of limits, total counts, or how to fetch additional results if there are >10 matches.