Open agent harness CLI and MCP server for coding agents, isolated worktrees, and software-factory workflows
HAR MCP Server provides 24 tools with reasonable naming conventions (all start with har_), but has significant inconsistencies in description quality, parameter documentation, and schema completeness. Naming is verb-based and clear (har_launch_environment, har_run_stage, etc.), but descriptions vary widely in clarity and actionability. Many tools lack clear output schema documentation. Parameter descriptions exist but are often generic. Error handling is minimal, most tools lack guidance on what to do if failures occur. The server targets a specialized domain (harness orchestration) with complex state management, but the tool interface doesn't consistently guide LLM usage.
Register an external code review, production incident, or other line-unit event in the factory. Appends to the work unit's line log.
Install a verification plugin (bundled id, local path, npm package, or git URL) that registers stages in .har/stages.json.
Append a related external link (PR, mirrored issue, alternate tracker) to an existing work unit.
Mark an agent slot as complete: commit work, run verification, release the slot. Returns verification result and commit hash. Idempotent after the first call (subsequent calls return the cached result); safe on uninitialized slots.
Start or re-attach Mission Control (the observability dashboard). Returns the port and access URL. Docker is required; the response reports Docker availability and any blocking issues.
Create a new line bundle for a work unit: a collection of verifications and gates that the factory orchestrates. Returns the line identifier.
Output schemas are undocumented for all tools. The rubric requires 'Document the output schema. LLMs need to know what fields to expect so they can plan downstream tool calls.' No tool in the source shows explicit return type definitions, input_schema type information, or example responses.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 60 | 2025-06-18+ | v2 |
Read manifest, stack hints, available scripts, and harness stages for a repository.
Validate the harness contract: harness.env schema, stages.json, stage scripts exist and are executable, lifecycle stages resolve, verificationStages ids resolve, port lanes are coherent, slot registry entries point at existing worktrees. Returns pass/fail with actionable findings.
Return status for all lines in the work unit: gate pass/fail, per-station results, and timing.
Return the status of a line (verification bundle): gate pass/fail, per-station results, timing, and logs.
Return recent logs for a slot/process.
Inspect a specific completed run: verification result, timing, pass/fail, and per-stage output for debugging.
Return structured slot status for one agent or all slots (same source as har env status/--json). Call BEFORE har_launch_environment when a slot may already be in use — shows worktree path, dirty state, branch, and readiness.
Scaffold .har/ boilerplate for a coding agent to adapt. Docker is required (Mission Control container + harness infra); the response reports Docker availability.
Start a FRESH agent session from the main checkout HEAD (switch that checkout to main first for a new unrelated task). Run BEFORE editing any file: returns workDir — ALL edits go there, never the main checkout. Occupied slots always block: call har_get_status, then har_complete_environment or har_teardown_environment, then relaunch — unless the session is failed/starting, in which case pass resume=true (or use har_recover_environment).
List run artifacts (logs, reports, evidence files) available for download or inspection.
List completed runs: ordered most-recent-first with status (passed/failed), timing, and commit hash. Use with har_get_run to inspect individual runs.
Validate .har/ against the bundled templates: returns validation issues, drift, and a maintenance bundle report. Pass finalize=true (optionally with summary) to record a completed manual adaptation in .har/manifest.json.
Readiness gate before launch: checks ports, foreign PM2, Docker conflicts, occupied slot, and untracked paths that will be missing from a session worktree. Returns canLaunch with actionable blockers and warnings. Call before har_launch_environment.
Resume a failed or partial agent launch without replacing the worktree. Alias for har_launch_environment with resume=true.
Run a factory gate: the bundled verification stages for a line in sequence. Returns pass/fail and per-stage output.
Run one generic harness stage by id or kind.
Run the project verification pipeline for an agent slot. Returns status, timing, and failed-step output (passing steps omit logs; stdout is not the raw JSON dump).
Stop and tear down a running agent slot: kill processes, remove the worktree, and release the slot for the next occupancy. Idempotent — safe to call on free/unknown slots.
Parameter descriptions lack clarity and actionability. Tools like har_run_stage have minimal guidance on 'args' parameter (just 'array of strings'), and tools like har_maintain do not explain what 'validation issues' or 'drift' mean. Many descriptions fail to meet the '10-1024 character' guideline and lack specificity.
No error handling or recovery guidance. Tools like har_launch_environment have complex preconditions ('Occupied slots always block') documented in the description, but no structured error responses or guidance on what to do if launch fails. Pattern:recovery-guide requires 'Error responses must tell the LLM what to do next.'
Missing tool annotations. The rubric rewards 'tool annotations (readOnlyHint/destructiveHint/idempotentHint)' in Spec Alignment. har_teardown_environment is marked DESTRUCTIVE and har_complete_environment is IRREVERSIBLE in the metadata, but there is no evidence these annotations are embedded in the tool definitions themselves for LLM visibility.
Complex parameter relationships and mutual exclusivity not documented. har_launch_environment has parameters like 'resume' that alter the semantics of 'repo' and 'agentId', but no parameter description explains that resume=true makes other parameters like 'worktree', 'claude', 'workUnitId', 'source', 'sourceUrl', 'title' inapplicable or have different effects.
No pagination documented for tools returning lists. har_list_runs, har_list_artifacts, and har_get_all_line_statuses do not specify limit/offset/cursor parameters, and descriptions do not mention how many results are returned or if results are capped. Rubric requires 'Tools returning lists should accept page/offset and limit parameters and return a total count or next_cursor.'
Descriptions fail to answer 'When should the LLM call this instead of a similar tool?' har_launch_environment, har_recover_environment, and har_preflight_environment have overlapping purposes but no clear distinction in descriptions. LLMs will waste reasoning cycles deciding between them.