On-demand micro-mutation sandbox for AI test verification — maps holes in unit tests by running isolated mutation testing via the Model Context Protocol
Three tools with strong naming (verb-first: audit_, triage_, estimate_) and comprehensive parameter schemas. audit_code_resilience has exceptional detail (lineScope, mutatorDenylist, baseline, enrich, diffBase, prebuildCommand, concurrency, dryRun, incremental). Descriptions are substantive (194 - 250 chars, within baseline p90=392). However, output schemas are NOT documented, the response structure for each tool is missing, forcing LLMs to infer what fields to expect. triage_test_coverage and estimate_audit descriptions lack specificity about return structure. No tool annotations (readOnlyHint, destructiveHint, idempotentHint) despite audit_code_resilience being read-only and estimate_audit being idempotent.
Runs on-demand, sandbox-isolated mutation testing against a single source file to identify gaps in unit test coverage. Chaos-MCP generates mutants (logical faults like changing `>` to `>=`) and checks whether the local test suite catches them. Surviving mutants indicate test coverage holes. Supports TypeScript/JavaScript (StrykerJS), Python (cosmic-ray), Rust (cargo-mutants), and PHP (Infection). PATHS: `filePath` is resolved against the SERVER's working directory (or given absolute). The `target` in the result is relative to the audited file's own WORKSPACE, which differs whenever the file sits in a monorepo package or another root — `packages/api/src/math.ts` comes back as `src/math.ts`. When the two differ the result carries `workspace` (an absolute path) and `join(workspace, target)` is a `filePath` you can pass straight back; when it is absent, `target` already is one. A `file` from triage_test_coverage is always a valid `filePath` as-is.
Estimates the time and resource cost of auditing a file before running the full mutation test. Useful for understanding whether a file is too large or complex to audit within a given time budget.
Scans a workspace to identify which source files lack test coverage, returning a ranked list of audit candidates. Useful for prioritizing which files to audit with audit_code_resilience. Returns files grouped by coverage status (untested, low coverage, etc.) with metadata to guide the audit.
Output schemas not documented. audit_code_resilience returns mutants, survivors, noCoverage, workspace, target, but the response structure is invisible to LLMs. triage_test_coverage and estimate_audit lack any return type specification.
Tool annotations missing. audit_code_resilience is read-only (no state mutation) but lacks readOnlyHint. estimate_audit is idempotent but lacks idempotentHint. These hints guide agent retry logic and safety reasoning.
triage_test_coverage description (58 chars) is too brief. It does not explain what 'coverage status' means (untested vs low vs high), what metadata is returned, or how to use results to prioritize audits. Baseline: descriptions should be 50 - 200 chars with actionable context.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | D | 59 | 2025-06-18+ | v2 |
estimate_audit description (95 chars) lacks clarity on what 'time and resource cost' means. Does it return milliseconds, complexity score, or a human-readable estimate? No guidance on when to call it vs audit_code_resilience directly.
Error handling guidance missing. audit_code_resilience can fail if the file is not auditable, test suite fails, or timeout is exceeded. No recovery hints (e.g., 'Try increasing timeoutMs' or 'Ensure test suite runs locally'). Baseline: errors should guide LLM to next action.