Durable execution and named concurrency queues for coding agents across local, SSH, and Slurm
Awaitless demonstrates strong naming conventions (verb-noun patterns: run, submit_job, wait_for_job, get_job_status, get_job_logs) and comprehensive parameter schemas with proper types and descriptions. Tool descriptions are detailed and explain WHEN to use each tool vs. alternatives, a pattern-aligned strength. However, output schemas are not documented in the visible code, and error handling guidance is absent from descriptions. The 'run' tool description is exceptionally long (500+ chars), violating the 10-1024 char guideline. Parameter descriptions are generally good (60-150 chars), but lack explicit format constraints (e.g., timeout_seconds lacks min/max bounds). No tool annotations (readOnlyHint, destructiveHint, idempotentHint) are present despite clear risk classifications (WRITE vs READ_ONLY).
Get bounded logs from a job for failure diagnostics.
Get the current status of a job without waiting.
Default tool for starting one non-interactive command of uncertain duration. Every command is a durable Job from launch. Commands that finish within the inline timeout return their ordinary bounded result. Longer or queued work returns a detached Job handle without cancelling the workload. Omit queue to use an operator-configured default queue for the selected target. Do not use this tool for explicit fire-and-forget or batch fan-out; use submit_job. Do not use it when an MCP Tasks handle is explicitly required; use run_job. Do not use it to resume an existing job_id; wait for or inspect that job instead. If a detached handle is returned, keep its job_id and call wait_for_job once when the result is needed rather than polling status.
MCP Tasks compatibility entry point for explicitly creating a Task handle. A client declaring io.modelcontextprotocol/tasks receives a Task handle immediately. Older clients block and receive the ordinary final tool result. The stable client_request_id makes a lost creation response safe to retry. Do not choose this for ordinary command execution: use run. Do not choose it for generic asynchronous submission or fan-out: use submit_job. Only retry a lost Task creation with the same client_request_id and
Output schemas not documented. Tool descriptions explain inputs but never specify what fields/structure the response contains. LLMs cannot plan downstream calls or extract required data without knowing response structure.
'run' tool description is 500+ characters, violating the 10-1024 guideline. Excessive length wastes tokens and buries key decision criteria. Should be condensed to ~150 chars with dependency hints moved to parameter descriptions.
Numeric parameters (timeout_seconds, stall_timeout_seconds, inline_timeout_seconds, tail_lines, max_bytes) lack min/max bounds in schema or description. LLMs may pass absurd values (e.g., timeout_seconds: 999999999) that break execution or cause timeouts.
Inferred effective spec: 2026-07-28+.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | D | 59 | 2026-07-28+ | v2 |
Explicitly submit one asynchronous or fan-out job without waiting. Omitted backend and host values use the Awaitless configuration defaults. Reuse client_request_id only when retrying the same logical submission; an identical retry returns the original job and a conflicting retry is rejected. A named queue provides FIFO, non-preemptive admission for local or SSH work. Slurm options may contain account, constraint, cpus_per_task, gres, mem, nodes, ntasks, partition, qos, or time. Cluster config supplies defaults. Do not use this as the default for a single command with uncertain duration; use run. Do not use it for MCP Tasks creation; use run_job. Do not resubmit merely because a client disconnected or a wait timed out: keep the original job_id, or retry the identical logical submission with the same client_request_id if the creation response was lost.
Wait for multiple submitted jobs to complete using a durable cursor.
Wait for a submitted job to complete and return its result.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) despite clear risk classifications. Tools marked WRITE (run, submit_job, run_job) should declare destructiveHint=true; READ_ONLY tools should declare readOnlyHint=true. Annotations enable safer agent planning.
Error handling guidance absent. Descriptions do not explain what errors can occur, whether they are retryable, or what the LLM should do next. E.g., 'job_id not found' should suggest 'Check job_id or call get_job_status() to verify it exists.'