MCP server and runner for autonomous PRD execution with Codex or Claude. Git worktree isolation, progress tracking, auto-merge.
ralph-mcp demonstrates solid definition quality with comprehensive parameter schemas, mostly clear descriptions, and good naming discipline. All 15 tools have explicit input schemas with typed parameters and descriptions. Tool names follow verb_noun conventions (ralph_start, ralph_status, ralph_update, etc.). However, several tools have overly complex schemas with nested objects and conditional dependencies that lack explicit documentation. Output schemas are not documented. Error handling guidance is minimal. The server shows strong foundational quality but lacks the refinement and LLM optimization of A-grade tools.
Start multiple PRD executions in batch.
Claim a ready PRD execution for processing by an agent.
Diagnose health and configuration of Ralph MCP server and runner.
Get detailed status of a single PRD execution including all user stories.
Merge a completed PRD to main branch with conflict resolution strategy.
Manage merge queue for PRD executions.
Reset stagnation counters for a PRD execution.
Output schemas not documented. No visibility into what each tool returns (e.g., ralph_status should document the structure of execution summaries, ralph_get should document the detailed PRD execution object). LLMs cannot plan downstream tool chains or extract needed fields without knowing return structure.
ralph_update has extremely complex nested input schema (acEvidence, hardGates, scopeExplanation objects) with undocumented conditional dependencies. Description does not explain when each nested parameter is required or how they interact. LLMs will struggle to construct valid inputs.
Descriptions for ralph_get, ralph_claim_ready, ralph_reset_stagnation, and ralph_merge_queue are terse (under 80 chars) and lack actionable context on when/why to call these tools. Missing guidance on typical workflows and prerequisites.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | B | 70 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 37 | - | v1 |
Retry a failed PRD execution.
Set agent ID for external agent task tracking.
Set runner concurrency limit.
Gracefully shutdown Ralph MCP server.
Start PRD execution. Parses PRD file, creates worktree, and stores ready/pending state for the Runner. Manual agent prompts are returned only when RALPH_AUTO_RUNNER=false. Dependencies must be merged before dependents start unless ignored.
View all PRD execution status. Replaces manual TaskOutput queries. Shows progress, status, and summary.
Stop PRD execution. Optionally clean up worktree.
Update User Story status. Called by subagent after completing a story.
No error handling guidance in any tool description. When a tool fails, LLM has no hint whether to retry, ask user, or abort. E.g., what happens if ralph_merge encounters a conflict? What if ralph_start hits a missing PRD file?
Tool descriptions mention side effects (e.g., 'Optionally clean up worktree', 'Auto add to merge queue') but do not make idempotency guarantees explicit. Agents retrying these tools risk duplicate work or resource leaks.
Parameters like 'onConflict' enum values ('auto_theirs', 'auto_ours', 'notify', 'agent') lack descriptions explaining what each strategy does and when an agent should choose each. LLM must guess.
ralph_status returns statuses like 'pending', 'ready', 'starting', 'running', 'interrupted', 'completed', 'failed', 'stopped', 'merging', 'merged' but state machine transitions and valid state progressions are not documented. LLM cannot reason about which transitions are legal.