Code intelligence MCP server for AI coding agents with impact analysis, semantic search, dependency graphs, and local release gates
The flyto-indexer server presents a sophisticated code-intelligence tool suite with strong descriptions and well-structured schemas. Most tools have clear, actionable descriptions (50-300 chars) and properly typed input schemas with parameter annotations. However, there are gaps in output documentation, error handling guidance, and some tools lack depth in parameter descriptions. The server demonstrates above-average quality overall but falls short of A+ (90+) due to missing recovery guides and incomplete response shaping documentation.
Comprehensive code quality audit. Starts with a health score (0-100), then automatically expands any dimension scoring below 80 with detailed findings: - Security < 80 → shows hardcoded secrets, injection risks - Complexity < 80 → shows complex functions, duplicates - Dead code < 80 → shows unreferenced symbols - Coverage < 80 → shows untested high-impact code Always includes git hotspots (high-churn + complex files) and stale symbols (heavily referenced but not modified in 180+ days), plus a bounded, noise-filtered local Git evidence portfolio and evidence-linked verdict. Use 'focus' to force expansion of a specific dimension regardless of score. focus='research_priority' answers a different question from the rest of the audit: not 'is this codebase healthy' but 'which code paths are worth a security researcher's next hour'. It runs a full taint scan and returns one ranked short list with the reasons attached, so it is opt-in rather than automatic.
Run impact analysis on multiple symbols at once. More efficient than calling impact_analysis repeatedly. Returns per-symbol breakdown and deduplicated affected list.
Detect file changes since last index and optionally clear caches. dry_run=true (default): only report which files changed. dry_run=false: clear all caches (must run 'python index_all.py' after). auto_reindex=true: detect changes AND perform live incremental reindex in-process. Returns: changed files grouped by type (modified/added/deleted) and project.
Track cross-project API usage. When a function/class in one project changes, find all other projects that need to be updated. Use this before changing shared APIs (e.g. a function in flyto-core used by flyto-pro and flyto-cloud). Returns: list of cross-project references, affected projects, and risk level (low/medium/high).
Output schemas not documented. Tools return complex structured data (lists of references, blast-radius summaries, git evidence portfolios) but the response schema is not explicitly defined in source. LLMs cannot plan downstream tool usage without knowing response structure.
Error handling guidance missing. No description of what errors each tool can return, what causes them, or how to recover. E.g., impact_analysis doesn't document what happens if symbol_id is invalid or if the project hasn't been indexed yet.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-21 | B | 73 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 26 | - | v1 |
Get the dependency graph for a file, symbol, or entire project. Shows what a module imports (dependencies) and what imports it (dependents). Use direction='imports' to see what a file depends on, 'dependents' to see what depends on it, 'both' for full picture. Returns: lists of import and dependent relationships with file paths and dependency types.
Preview the impact of editing a symbol before making changes. Shows exact symbol identity, same-name ambiguity, call sites, unresolved dynamic references, required update/test sites, and risk. Use before a rename, move, delete, or signature change.
Find unreferenced functions, classes, and components (dead code). These symbols are never imported or called by any other code and can likely be removed. Automatically excludes entry points, lifecycle hooks, private methods, and test files. Returns: list of dead symbols sorted by line count (largest first), with total dead lines.
Find all places that call or import a specific symbol. Use this BEFORE modifying a function/component to understand who depends on it. Uses pre-computed reverse index for fast, accurate results. Each reference includes confidence level: high (from dependency analysis), medium (from imports), low (from content regex). Returns: list of callers with file path, line number, and confidence, grouped by project.
Analyze what breaks if you change something. Two modes: 1. Symbol mode (target): finds all references, blast radius, cross-project impact, and related test files — all in one call. 2. Diff mode (mode): analyzes uncommitted/staged/committed changes, maps affected symbols, and finds their test files. Auto-enriches with: - Cross-project impact (if multiple projects indexed) - Test file mapping for affected code - Semantic preflight for rename/move/delete/signature changes - Call path tracing (entry points → target) - Relevance-scored references (recency, confidence, proximity) Diff mode also returns a bounded, noise-filtered local Git evidence portfolio and a concise verdict whose claims link to source receipts.
Analyze the blast radius of modifying a symbol. Use this to assess risk BEFORE making changes to shared code. Returns: count of affected locations, list of affected symbols with paths, and a risk assessment (safe / moderate / high risk) with suggestions.
Parse git diff, match changed hunks to indexed symbols, classify each change (signature_change, body_change, rename), and run impact analysis.
List all indexed projects with statistics. Use this FIRST to discover available projects and their sizes. Returns: project names, file counts, symbol counts, and breakdown by symbol type (function/class/component/etc).
Find code by keyword or natural language. Runs BM25 keyword search AND semantic search (TF-IDF with learned concept expansion) simultaneously, merges results, and auto-enriches top hits with: - Callers: who calls this symbol (top 5) - File siblings: other symbols in the same file Use this for ALL code search needs — no need to pick between search modes.
Interrogate, plan, gate-check, validate, or learn from code changes. Five actions: 1. grill: Build a persistent evidence-backed decision tree before planning. Repository-owned facts are resolved from the code index; human decisions are ordered by confidence and value-of-information, then asked one at a time with a recommendation and bounded adversarial review. Frozen v2 contracts include content-addressed evidence snapshots, selective-reopen metadata, acceptance criteria, ADR Markdown, and a machine-readable audit artifact. Critical unresolved decisions and contradictions fail closed. Operations: start, answer, status, freeze, discard. 2. plan: Build one change contract: risk, target-scoped agent instructions, requirement-to-plan traceability, and co-change evidence. A frozen grill_session_id also attaches the decision contract. 3. gate: Check if you can proceed to the next phase. Server-side enforcement blocks skipping required gates. If pass=false, execute every required_actions item, update current_state with the exact requested keys, and immediately re-run the same gate until pass=true. Do not enter the blocked phase or edit while the gate is false; ask the user only for unavailable authorization, required input, or an external state change. 4. validate: Run ruff + pytest after making changes. When task_contract is provided, also verify instruction/spec freshness, requirement coverage, declared proof results, and decision-to-diff conformance. The result is recorded in a privacy-preserving local outcome store to calibrate future confidence. Optional attested external proof receipts can close browser, race, container, security, integration, and deployment requirements without embedding those runtimes. Auto-attaches untested change analysis if validation fails. 5. feedback: Record, summarize, or resolve local AI-development problems such as false positives, missing context, framework gaps, slow scans, and bad recommendations. Feedback never stores prompts or source code and cannot automatically weaken policy. Workflow: grill → freeze → plan → gate → edit → validate → feedback
Parameter dependency documentation incomplete. The 'task' tool's grill/plan/gate/validate/feedback workflow has complex state dependencies (e.g., 'gate requires task_contract from plan'), but these are buried in description text rather than formalized in schema or clear error messages.
Pagination/result limits not enforced for all list-returning tools. Tools like 'search' and 'impact' can return unbounded lists (affected symbols, call sites, evidence items). Descriptions mention results but don't explicitly state limits or pagination strategy.
check_and_reindex has a risky default: dry_run=true (report only) vs dry_run=false (clear caches). The description warns that cache-clearing requires manual 'python index_all.py' afterward, but there's no confirmation step or idempotency marker. An LLM could accidentally trigger cache invalidation expecting non-destructive behavior.