Scores every function on complexity times uncovered risk, ranks the worst, and blocks commits that add more. Provides MCP tools for code quality analysis via CRAP score (Cyclomatic Complexity Risk Analysis and Predictions).
crapkit MCP server exhibits strong definition quality with detailed descriptions and well-structured tool schemas. All 12 tools have meaningful, domain-specific descriptions (average ~250 chars, well above baseline 194). Parameters are consistently typed and described. However, output schemas are not explicitly documented in the visible source code, and error handling guidance is minimal. Tool names follow verb_noun convention and are appropriately specific to code quality domain (get_next_item, list_worklist, get_function_brief, etc.). The server demonstrates sophisticated parameter constraints (enums for remedies, exclusive parameters with clear documentation). Missing elements: (1) explicit output schema documentation for tool results, (2) error recovery guidance (what to do when a tool fails), (3) tool annotations (readOnlyHint/destructiveHint). The domain complexity (cyclomatic complexity analysis, mutation testing, coverage metrics) is well-represented in parameter descriptions, but LLMs would benefit from clearer guidance on which tool to call in which workflows.
Returns the function's current score breakdown (ccn, coverage, crap, remedy) and its history under the changes the repo makes. Pass name (long name or brief from get_next_item), and optionally path when the function is ambiguous.
Returns the function's history: every run's score, remedies reached, and ratchet marks; walks the git history in the churn window and names the commits that changed the function.
Returns the result of one mutation test: what changed, whether it was caught (killed) or escaped (lived), and the test evidence gathered during its kill attempt.
Returns the source for one mutation: the original and the changed lines, plus enough context to see what was changed and why. Each row is identical to what an editor would show (where a mutation changed lines inside a function, get_mutation_source numbers them 1-based inclusive).
Returns the next function to fix as a work packet, the newest trusted run's worst row by crap. Use it to start a fix, list_worklist to survey the same run by risk, and get_function_brief once a function is chosen. It runs no tests, and empty true means the queue is spent, not the work, so read reasons. Filters cut before top counts: top 3 with exclude ["tests/"] returns the three worst rows outside tests, and an unknown scope name is a config error.
Output schemas not explicitly documented in source code. Tools return results via subprocess CLI invocation, but response structure is not formally defined for LLM consumption. get_next_item returns 'items', list_worklist returns a 'list', but field names, types, and nesting are inferred from CLI behavior rather than schema. This violates the 'document output schema' requirement.
No error handling guidance. Tool descriptions state prerequisites (e.g., 'runs no tests', 'empty true means the queue is spent') but do not document what errors LLMs should expect or how to recover. For example, score() and verify() are WRITE operations that can fail (missing config, git errors, test failures), but no error messages or recovery steps are documented.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 68 | 2025-06-18+ | v2 |
Tells the client the server is responsive.
Inventories the repo without running lanes: parses every source file and reports all functions declared, grouped by scope. The output carries no coverage, crap score, or gate status; use score when those matter.
Lists all open findings from the newest gate run (latest run only; no history), filtered by path and remedy and ranked by worst-first crap score. Returns all when no filters are set; empty true means no run or all filtered.
Lists all mutations in the newest mutation test run, filtered by path and remedy, and ranked by worst-first crap score. Empty true means no run or all filtered; when false, every row's source file is the artifact its mutation ran, which can differ from the checked-out tree.
Lists all functions in the newest trusted run, ranked by risk (complexity-weighted churn), filtered by exclusions and scope. Use it to survey work by risk when the run is measured, or to get an inventory when it is not. When measured, every row but ok reaches get_next_item; when only inventoried, only get_next_item's top item is usable.
Scores the repo: runs every declared lane, ranks the worst functions by crap, and runs every declared mutation test if the gate passes or if mutations are pre-scored. Returns the run's full metadata, including status and any step that failed.
Verifies the repo: runs every lane and checks gate status without updating ratchet marks, and skips mutations. Useful for checking a change without committing.
Tool annotations missing. score() and verify() are WRITE operations; list_* and get_* are READ_ONLY. The source code notes 'Risk: READ_ONLY' and 'Risk: WRITE' but these are not emitted as MCP tool annotations (readOnlyHint, destructiveHint, idempotentHint per spec 2026-07-28). LLMs cannot detect which tools are safe to retry vs. which have irreversible consequences.
health() tool has a trivial description ('Tells the client the server is responsive.'). This is a supervision/liveness check that does not need to be exposed as an LLM-callable tool; it should be handled by the MCP infrastructure (ping heartbeat), not invoked by agents.
Inconsistent parameter defaults. Most tools default 'repo' to 'the repo the server was started in', but no explicit default value is shown in the schema. LLMs cannot infer whether a parameter is optional and what the default is. Schema should include 'default' field or description should say '(optional, default: current directory)'.
Parameter naming inconsistency: get_function_brief and get_function_history accept 'name' as a complex union (bare identifier OR full long_name), but the description is dense and hard to parse ('the bare identifier (classify, or route for a Rust `route cmd : & Cmd`) or the whole long_name next_item printed'). This violates the naming clarity rule, the parameter should be clearly named and the description simplified.