LLM-native code toolkit with 44 MCP tools for AST-based code analysis, symbol extraction, refactoring, testing, and impact analysis. Uses tree-sitter for multi-language parsing with Rust backend.
This MCP server exhibits critical definition quality gaps across nearly all 45 tools. Evidence: (1) Tool definitions are inferred from test files (malong/tests/test-tool-registry.js) rather than explicitly visible in registration code, tools like 'read_symbol', 'write_symbol', 'edit_batch' have descriptions and Risk labels listed in test assertions, but the actual tool schema definitions are not provided in the source snippet. Per HARD SCORING RULE, inferred tools cap at 50 per-tool. (2) Of the few tools with visible schemas (health, reindex, code_search, read_outline), most lack complete parameter descriptions. For example, 'health' shows only {'type':'object','properties':{'action':{'type':'string','description':'check action'}}}, vague parameter description ('check action' is not actionable). (3) No output schemas are documented anywhere in the provided source. The absence of documented return types violates baseline requirement: '100% of A+ tools have documented return types.' (4) 40+ tools listed with only brief one-line descriptions (e.g., 'Read symbol definitions and outlines from source files', 'Execute tests and analyze test results'), many fall below the 50-character minimum useful description length and fail to state WHAT, WHEN to use, and what it RETURNS. (5) No evidence of error handling patterns, recovery guidance, or actionable error messages. (6) Tool names use inconsistent verb patterns: some follow verb_noun (read_symbol, write_file), others are noun-heavy (repo_map, config_drift, workflow_closure). Names like 'sweep_dead_code' vs 'dead_code_sweeper' show inconsistency. (7) No parameter validation rules, enums, ranges, or defaults documented. For destructive tools like delete_files, execute_shell, git_operations, no dry-run or confirmation patterns are evident. (8) No evidence of IDs/references matching between tool outputs and inputs (e.g., does rename_symbol return a symbol_id that other tools accept?). (9) Security-critical tools (execute_shell, delete_files, git_operations) have no visible permission gates, scope declarations, or audit trail patterns.
Extract and track TODO and FIXME comments
Trace call chains and dependencies between symbols
Assess code quality metrics and issues
Full-text search across indexed codebase
Detect configuration inconsistencies and drift
Comprehensive dead code detection and removal
Delete files or directories
Project dependency graph analysis
CRITICAL: 42 of 45 tools lack visible input schemas. Schemas are either missing or inferred from test assertions rather than explicit tool registration code. Per HARD SCORING RULE: tools with NO input schema must score schema=0.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 32 | <=2025-11-25 | v2 |
Enforce dependency access control policies
Batch edit multiple files with atomic transactions
Detect concurrent modification conflicts during edits
Transactional edit with rollback capability
Detect missing exception handling in code
Execute shell commands in workspace
List and manage tool feedback
Find test files and test cases related to code
Fix and optimize import statements
Git version control operations
Detect code pattern violations and anti-patterns
Health check and diagnostics
Analyze impact of code changes on callers and dependencies
Deep inspect of code structure and metadata
List files in directory with filtering
Synchronize and validate mock implementations
Check naming conventions and consistency across codebase
Monitor and enforce output token budgets
Read file contents with optional range selection
Extract file outline with token savings estimate
Read symbol definitions and outlines from source files
Read multiple symbols in batch
Reindex workspace for code search and symbol extraction
Rename definition line (function/class declaration)
Rename symbols with automatic refactoring across codebase
Generate repository structure map and symbol index
Validate code against sandbox execution constraints
Identify and remove unused imports, functions, and orphan files
Execute tests and analyze test results
Trace symbol references and usages throughout codebase
TypeScript type checking without config
TypeScript compiler specification validation
Trace variable references and bindings
Verify code transformations and edit pipelines
Analyze workflow execution closure and dependencies
Write or create files
Write or modify symbol definitions in source files
CRITICAL: Descriptions for 40+ tools are extremely brief (one-line, under 50 characters) and fail to convey WHAT, WHEN to use, RETURN value, or dependencies. Examples: 'Deep inspect of code structure and metadata', 'Extract and track TODO and FIXME comments', 'Execute shell commands in workspace'.
CRITICAL: No output schemas are documented. LLMs cannot plan downstream tool calls or extract relevant data without knowing what fields to expect. Pattern requirement: '100% of A+ tools have documented return types'.
HIGH: Destructive/sensitive tools (delete_files, execute_shell, git_operations, write_file, edit_batch, sweep_dead_code) show no visible permission gates, scope declarations, audit trail patterns, or confirmation/dry-run mechanisms. No evidence of 'permission-gate' or 'confirmation-request' patterns.
HIGH: No error handling guidance visible. Tools lack recovery instructions, categorization of retryable vs fatal errors, or actionable error messages. Example: if 'execute_shell' fails, what should the LLM do next? No guidance provided.
MEDIUM: Tool names lack consistency in verb-noun structure. Examples: 'sweep_dead_code' vs 'dead_code_sweeper' (inconsistent word order), 'inspect' (no verb prefix), 'active_todos' (noun-only), 'guard_patterns' (ambiguous).
MEDIUM: Three tools with visible schemas (reindex, code_search, read_outline) have incomplete parameter descriptions. 'action' in health tool described as 'check action' (not actionable). Parameter descriptions must state expected format, range, and valid values.
MEDIUM: No evidence of tool chaining support. Unknown if output fields (IDs, references) from one tool are compatible with input parameters of downstream tools. Per pattern: 'Return IDs and references that downstream tools accept.'
LOW: Source code snippet does not include the actual tool handler implementations or registration logic. Tools are inferred from test assertions and directory structure. Per HARD SCORING RULE: inferred tools cap at 50 per-tool overall.