MCP server for autonomous PRD execution with Claude Code. Git worktree isolation, progress tracking, auto-merge.
ralph-mcp provides 12 well-defined tools for PRD execution management with explicit tool registration, comprehensive input schemas, and tool annotations. However, output schemas are not documented, error handling guidance is minimal, and several parameter descriptions lack specificity about formats and constraints. Tool naming is clear and verb-driven (start, status, get, update, stop, merge, retry, doctor). Most tools have descriptions in the 60-150 character range, which is solid. All 12 tools have explicit input schemas with types and required fields. The main gaps are: (1) no documented return schemas or output field descriptions, (2) limited error handling patterns, tools don't explain what the LLM should do if they fail, (3) some parameters lack format/constraint details (e.g., 'branch' format, allowed characters), and (4) missing context about idempotency and side effects beyond the annotations. The tool annotations (readOnlyHint, destructiveHint, idempotentHint) are present and correct, which is a strength. Composition is good, each tool has a single responsibility. Overall, this is a solid foundation that would benefit from documented output schemas and more actionable error messages.
Start multiple PRDs with dependency resolution. Parses all PRDs, creates worktrees, runs pnpm install serially (avoids store lock), and returns agent prompts in execution order.
Run diagnostics to verify MCP server setup and git configuration.
Get detailed status of a single PRD execution including all user stories.
Merge completed PRD to main and clean up worktree. MCP executes directly without Claude context.
Manage merge queue. Default serial merge to avoid conflicts.
Reset stagnation counters for an execution. Use this after manual intervention to allow the agent to continue.
No documented output/return schemas. Tools list descriptions of what they do, but agents cannot see what fields will be returned (branch names, status strings, story IDs, etc.). This forces LLMs to guess at response structure and breaks tool composition, agents cannot reliably extract the branch name from ralph_start to pass to ralph_update.
Limited error recovery guidance. Tool descriptions do not explain what the LLM should do if a call fails. For example, ralph_start says 'Fails if dependencies are not satisfied unless ignoreDependencies is true', but what error code does it return? Should the LLM retry, call ralph_doctor, or ask the user? No guidance.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 69 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 33 | - | v1 |
Retry a failed/interrupted PRD execution. Supports resuming from breakpoint with WIP handling.
Record the Claude Task agent ID for an execution. Called after starting a Task agent.
Start PRD execution. Parses PRD file, creates worktree, stores state, and returns agent prompt for auto-start. Fails if dependencies are not satisfied unless ignoreDependencies is true.
View all PRD execution status. Replaces manual TaskOutput queries. Shows progress, status, and summary.
Stop PRD execution. Optionally clean up worktree.
Update User Story status. Called by subagent after completing a story.
Parameter 'branch' used across 8 tools but lacks format specification. Descriptions say 'Branch name (e.g., ralph/task1-agent)' but do not enforce the format, are forward slashes required? What characters are allowed? Can underscores be used instead of hyphens? LLMs will guess, potentially passing invalid values.
ralph_merge_queue 'action' parameter has enum values ['list', 'add', 'remove', 'process'] but descriptions are missing for each action. What does 'process' do? Does it merge queued branches serially? Partially? With conflict resolution? Ambiguous.
No idempotency guarantees stated in descriptions. ralph_update sets story status, if called twice with the same values, does it idempotently succeed or error? ralph_retry restarts execution, is it safe to call twice? Agents need to know whether to retry on failure or abort.
Confirmation/dry-run pattern missing for destructive operations. ralph_merge and ralph_stop can permanently delete worktrees and execution records, but no confirm_before_execute pattern or force flag with validation. Agents could accidentally destroy work.
ralph_start 'prdPath' parameter lacks format guidance. Should it be relative to projectRoot? Absolute? Must it end in .md? What happens if the file doesn't exist? Agents cannot reliably construct valid paths.
No per-item error reporting for batch operations. ralph_batch_start processes multiple PRDs but does not document whether it returns per-PRD success/failure or a blanket error if any PRD fails. If one PRD has missing dependencies, do all batch starts fail?