Root-cause analysis engine for git repos: classify feature vs bugfix, find the introducing commit (suspect set) or the blast radius. RCA (root-cause analysis) engine for git repos. Recommended agent workflow: 1. analyze() - full overview in one call (classify + suspects/blast-radius + risk + test impact) 2. find_suspects() - drill into which commits introduced the bug 3. get_evolution() - line-by-line history of the exact buggy lines 4. get_intent() - what the author was trying to do in the suspect commit 5. check_completeness()- are there other call sites the fix missed? 6. verify_fix() - check proposed diff pre-commit; iterate until verdict='complete' (then add a test if risk_level flags one) 7. get_risk_score() - QA gate score (use with --fail-on in CI). For a stack trace with no fix in hand, start with from_trace() instead of analyze(). All tools are read-only; they never modify the repo or create commits.
culprit exposes 11 well-intentioned RCA tools via FastMCP with generally clear naming and solid descriptions. However, critical gaps exist: (1) Input schemas are inferred from source code comments rather than formally visible in tool registration; (2) Output schemas are documented in descriptions but not in structured format; (3) Error handling guidance is minimal, tools return structured results but lack 'what to do next' guidance for failures; (4) No parameter validation constraints (enums, ranges, patterns) are visible in the schema definitions. The server demonstrates domain expertise (RCA patterns are sophisticated and well-composed) but falls short of production-grade tool definition quality. Average per-tool score: 68.
Full RCA in one call: classify -> suspects (bugfix) or blast-radius (feature) -> risk score -> test impact. Returns the complete structured result. Use the individual tools to drill into specific signals.
Is the fix complete? Find other references to changed symbols not touched by this fix. Returns: {symbols: [...], other_call_sites: {...}, untouched_count: int, adds_test: bool, is_revert: bool, notes: [...]}
Classify whether a change is a bugfix or a feature, with evidence. Returns: {verdict: "bugfix"|"feature"|"unknown", evidence: [...], signals: {...}}
Find the commits most likely to have introduced a bug. Pass trace_text (a stack trace / crash log) to run RCA from a runtime error with no diff needed. Otherwise diffs base..head to find suspects. Returns: {suspects: [{hash, short, author, date, subject, pr_number, weight, lines}], origin_on_branch: bool, notes: [...]}
RCA from a stack trace or crash log; no diff or PR needed. Parses the stack trace, blames the crashing lines in git history, and returns the suspect set. Works for Python, JavaScript, Java, and Go stack traces. Returns: {suspects: [...], frames: [{file, line, func}], skipped_frames: [...], notes: [...]}
Input and output schemas are documented in natural language descriptions but not formally registered in tool definitions. No visible JSON Schema type definitions, constraints, or enums in the tool registration code.
Error handling provides no recovery guidance. Tools return structured results but lack actionable next steps for failures. E.g., 'no suspects found' should suggest 'try with a different base/head ref' or 'check trace format'.
Output schemas documented in descriptions lack formal specification. LLM must infer structure of complex fields like 'factors' (get_risk_score), 'by_test' (get_test_impact), 'other_call_sites' (check_completeness). No pagination or result limits documented despite tools potentially returning large datasets.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 63 | 2026-07-28+ | v2 |
Map what a feature change affects: who imports the changed modules, covering tests, high-risk areas. Returns: {dependents: {...}, covering_tests: [...], high_risk: [...], notes: [...]}
``git log -L`` over a line range: every commit that touched those lines, oldest to newest, with per-step diffs. Returns: {steps: [{hash, short, author, date, subject, diff}], notes: [...]}
Commit body + the PR it came from (title, body, url) + linked issues (Fixes/Closes/Resolves #N). Returns: {body: str, pr: {number, title, body, url} | null, linked_issues: [...]}
QA risk score for a change: 0-100 with level (low/medium/high) and contributing factors. Combines test gap, fix completeness, hotspot recurrence, blast radius, and churn. Returns: {score: int, level: "low"|"medium"|"high", factors: [{name, detail, points}]}
Which existing tests should be run for this change. Walks the reverse-import graph from changed files to tests that cover them directly or transitively (up to 2 hops). Returns: {tests: [...], by_test: {test: [reasons]}, notes: [...]}
Check fix completeness against a raw unified diff before committing. Runs completeness + test-impact analysis on the proposed diff and returns a verdict. Iterate until verdict == "complete" (no untouched call sites). "complete" covers the root cause but does not imply a test exists - a complete but untested fix comes back at risk_level "medium" with a note, so check risk_level/notes and add the test before committing. Returns: {verdict: "complete"|"partial"|"risky", symbols_fixed: [...], untouched_references: [...], tests_to_run: [...], adds_test: bool, risk_level: "low"|"medium"|"high", notes: [...]}
Input parameters lack validation constraints. No enums for git refs, no ranges for PR numbers or line numbers, no format documentation for trace text or unified diffs. Parameters like 'base' and 'head' accept strings with no documented format (SHA, branch name, tag, range syntax?).
Some tool names are domain-jargon rather than action-verb patterns. 'get_evolution' (what is evolution to a non-git expert?), 'get_intent' (intent of what?), 'from_trace' (missing verb). Clearer names: 'trace_line_history', 'get_commit_context', 'analyze_stack_trace'.
Tools return multiple related fields (e.g., verdict + risk_level + notes in verify_fix) but no clear guidance on field presence or interdependencies. Verify_fix description hints that fields differ by verdict state but schema does not formally specify this.