Static-analysis-powered context compiler for LLMs. Selects the code that matters for a change and packs it into a token budget. Measures whether it works on your repo.
DiffContext MCP server has 4 tools with reasonable names and descriptions, but significant gaps in schema completeness, parameter documentation, and error handling. All tools are named with action verbs (compile_, find_, explain_, verify_) which is good. Descriptions are present and moderately detailed (100-300 chars each), exceeding the minimum but lacking LLM-optimized brevity. However, critical schema issues emerge: while input parameters are documented with types and descriptions, output schemas are entirely absent from the visible code. This is a major gap, agents cannot plan downstream operations without knowing what fields to expect. Parameter descriptions are adequate but could be more prescriptive about constraints. Error handling is not visible in the provided code. The tools themselves are well-scoped (each does one thing), but the MCP integration lacks modern patterns like tool annotations (readOnlyHint/destructiveHint), structured error responses, and recovery guidance.
Compile LLM-ready context for a change. Give it changed symbol IDs (e.g. ./src/auth.py:validate_jwt) or a git ref (e.g. HEAD~1), and it returns the callers, callees, and related functions the model needs to make the change safely — packed into max_tokens with a disclosure header showing what was dropped. Uses the gap cutoff by default (~6-9 symbols at ~4x precision vs top-20, costing ~30% recall) — MCP consumers are budget-sensitive, so the default is precise, not exhaustive. The total count of dropped symbols is always shown in the meta header. Optionally pass task_description (the bug report or issue text) to bias retrieval toward symbols relevant to the described problem — the one signal the graph alone can't provide.
Explain why symbols were included or dropped from context. Returns the included symbols (with scores and token costs) and the top dropped symbols (scored but cut by the token budget), so an agent can inspect or filter the selection. Dropped symbols are capped at 10 by default (gap cutoff) — the total count is always shown so nothing is hidden.
Find what breaks if you change a symbol. Returns the blast radius: direct callers, direct callees, and transitive impact, ranked by impact score. Default cap is 10 symbols — a reviewer brief must be short enough to read. Pass limit > 0 to override with a different cap, or limit=0 for all. The total count is always shown ("339 impacted, showing top 9") so nothing is hidden — the cap is a display choice, not information loss.
Verify retrieval quality on your repo using ground truth from git history. Runs the co-change benchmark: for each commit that changed multiple symbols, it picks one as the query and checks if DiffContext retrieved the others. Returns precision, recall, and F1 — the industry standard for retrieval evaluation. This is the HONEST eval: ground truth comes from what humans actually changed together, not from the graph (which would be circular).
Output schemas completely undocumented. No visible definition of what fields each tool returns. Agents cannot infer downstream tool calls or extract required data without explicit schema documentation.
No tool annotations present. Tools marked as read-only (all have Risk: READ_ONLY in metadata) but this is not expressed via readOnlyHint in the MCP schema, forcing agents to infer safety from descriptions alone.
Error handling strategy not visible. No documented recovery guidance, error classifications, or actionable error messages in the visible code. compile_context mentions 'disclosure header showing what was dropped' but no error contract is defined.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | D | 59 | 2026-07-28+ | v2 |
Parameter constraints underspecified. max_tokens accepts integers but no bounds are documented (minimum, maximum, default behavior if exceeded). meta accepts string enum values but not formally declared as enum in schema.
Descriptions exceed LLM-optimized range. compile_context is ~350 chars, exceeding the 10-1024 baseline sweet spot (p50=194). Long descriptions waste tokens and bury key selection signals.
Mutually exclusive parameters (changed_symbols vs git_ref) are documented in description text but not enforced in schema. Schema should express this constraint formally.