MCP server for project definition, verification, and status tracking with criterion-based testing and mutation checking
Rework exposes 5 tools with detailed docstrings and complex parameter schemas. Naming follows verb_noun convention (project_define, project_check, project_status, project_list, project_get). Descriptions are comprehensive (200-400+ chars) and explain WHEN to use each tool. However, parameter descriptions in the schema are minimal or absent, the JSON schema shows parameter names and types but lacks inline descriptions for most fields. Output schemas are not documented. Error handling guidance is missing. The server is STDIO-only, which caps protocol readiness at 50 regardless of definition quality.
Run FAIL_TO_PASS/PASS_TO_PASS verification for one criterion (Instructions §4). A criterion's test must pass a mutation check (§4.4) before its first real verification: pass mutation_target_file (the source file the test exercises) on that first call, and mutation_target_function (the specific function it calls) whenever the file has more than one function -- without it, the mutation may land in unrelated code the test never touches, producing a false 'not real verification' result for a fine test. If that function itself has multiple independent branches, pin mutation_target_line (1-indexed) to the exact line the test is meant to exercise. If the mutation check fails (couldn't even apply, or applied but the test didn't catch it), the tool won't guess who's at fault. Pass mutation_review_choice='human_review' to flag it for the user, or 'claude_review' to render a real verdict yourself right now -- in that case you MUST also pass claude_review_verdict ('trustworthy' or 'not_trustworthy') and claude_review_reasoning (a real, specific explanation, not a restatement of the verdict). 'trustworthy' reaches a distinct status, verified_complete_reviewed -- never plain verified_complete, so a judgment call always stays visibly different from an automated proof. Do not call this with mutation_review_choice='claude_review' as a label without actually reading the test and mutation result first; the reasoning you write is recorded on the criterion for anyone auditing this later.
Persist a project definition already interrogated and self-audited in conversation (Instructions §1-2). criteria_inputs items need: id, feature, given_when_then, q1_underspecified, q2_gameable, grounding (a quote/paraphrase of what the user actually said that this criterion traces back to -- Instructions §1.5), and optionally non_functional_risk, fail_to_pass_test, pass_to_pass_tests. Rejects (without saving) if criteria_inputs is empty, if two criteria share an id, or if a given_when_then still contains an unquantified adjective (secure, fast, accessible... -- Instructions §1.6). The plan should be comprehensive -- cover the real scope of the work -- not calibrated to any skill level.
Parameter descriptions missing or minimal in JSON schemas. Fields like 'changed_files', 'designated_test_paths', 'mutation_target_file' lack inline descriptions in the schema object, forcing LLMs to infer meaning from names alone.
Output schemas not documented. Tools return dict but no schema is provided describing the structure of returned fields (e.g., what does project_list() return? What fields are in each project summary?).
No error handling guidance. Tools do not document what errors can occur, whether they are retryable, or what the LLM should do next (e.g., if project_id not found, should it call project_list() first?).
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 69 | 2026-07-28+ | v2 |
Return a project's full stored definition -- goal, dod, reviewer_level, and every criterion in full: id, feature, given_when_then, grounding, self_audit, non_functional_risk, status, fallback_reason, and the complete verification record (including any claude_review reasoning). Use this when you (or a fresh session/second client) need to recall what a criterion actually requires -- project_status only reports which bucket each criterion id falls into, never the given_when_then text itself. Call project_list() first if you don't already know the project_id.
List every tracked project with a compact status summary each -- project_id, goal, reviewer_level, total_criteria, percent_complete, and each bucket's count (verified_complete, verified_complete_reviewed, fallback_complete, failed, in_progress, not_started). Use this to discover what project_ids exist before calling project_status or project_check, which both require an exact id you'd otherwise have no way to look up. Call project_status(project_id) for the full per-criterion breakdown of anything listed here.
Report verified-complete, verified-complete-reviewed, and fallback-complete counts, always separately (Instructions §5).
STDIO transport only. Server is not remotely accessible and cannot be used by hosted MCP clients. This is a hard architectural limitation.
Complex nested parameter structures (criteria_inputs in project_define) lack detailed field-level descriptions. The rubric requires descriptions for every parameter; nested object fields should also be documented.