The control side of heliograph: CLI and MCP server. Run commands on a machine you cannot SSH into.
heliograph demonstrates strong naming conventions and clear, contextual descriptions that guide LLM behavior well. All 7 tools follow verb_noun patterns (heliograph_estates, heliograph_send, heliograph_status, heliograph_logs, heliograph_read_log, heliograph_gaps, heliograph_doctor). Tool descriptions are detailed and educational, ranging from 100-380 characters with explicit guidance on dependencies and recovery paths. However, schema quality is inconsistent: input parameters are properly typed and described, but output schemas are entirely undocumented, LLMs have no visibility into response structure. Error handling guidance is present in descriptions but not formalized in error response definitions. Overall composition is sound: tools chain logically (send → status → read_log) and accept human-friendly identifiers (step names, estate names). The server handles state management well for a polling-based system.
Test the transport without changing anything. Publishes a noop request and verifies that the station can write a status and a log, then deletes both.
List the configured estates and the transport each uses. Call this first when you do not know which estate to act on.
Where a run stalled. Returns the intervals in the log's timestamp column, longest first, each attributed to the line BEFORE it, which is what was running. Do this before reading a long log: a hang and slow progress are indistinguishable without it. If every line carries the same timestamp the capture was buffered, and a buffered log cannot answer the question at all.
List the captured logs, newest first. Each name carries the request id that produced it, so the log for a run you sent is the one whose name matches the id heliograph_send returned. Use this to find a log; use heliograph_read_log to read one.
Read a captured log whole. Every line carries a UTC timestamp. Read all of it, including the parts that worked: a passing probe beside a failing one is the control that says what the failure means. A green exit means the probes that ran passed, not that the work happened.
Output schemas are completely undocumented. No tool documents return types, fields, or structure. LLMs cannot reason about response structure or plan downstream tool calls confidently.
Error handling is described in tool descriptions but not formalized. No error response schema, error codes, or categorization (retryable vs user-fixable vs fatal). Agents cannot self-correct on failures.
heliograph_send is destructive (WRITE) but has no confirmation step or dry-run mode. Agents cannot safely preview consequences before executing. Documentation warns about CONFIRM=yes but no tool-level protection exists.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 64 | 2026-07-28+ | v2 |
Publish a step for the station to run, and return immediately. This does NOT wait for the result: poll heliograph_status until it reports a terminal state, then read the log. The station decides whether to run it at all: a step that declares no mode is refused, and one that changes state needs CONFIRM=yes and a station started with --allow-actions.
What the station is doing now. State is one of starting, running, idle, cancelled, refused, stopped or undelivered. Five of those are terminal: idle, cancelled, refused, stopped and undelivered. `starting` and `running` mean the run is still going. An empty state means the station has published nothing yet, which usually means it has not been started. A state you do not recognise is neither finished nor alive: a newer station may publish one, so keep polling rather than giving up. `refused` is not a failure: it means the station would not run the step, and the reason names the flag that would permit it. `undelivered` means the run finished and the log is complete but the transport would not take it, so waiting for more is exactly wrong.
No tool annotations present (readOnlyHint, destructiveHint, idempotentHint). Only tool Risk labels visible in metadata; not reflected in MCP schema. Agents cannot determine safety/side-effect profile from protocol.
Pagination not supported. Tools like heliograph_logs and heliograph_read_log could return very large result sets. No limit, offset, or cursor parameters documented. Risk of context window overflow.