A tactical agent interface for deepagentsjs with multi-mode task orchestration (default, ralph, email, loop, expert modes)
OpenWork presents a mixed quality picture. The server provides 10 tools across two domains (Butler task creation and Docker operations), all with JSON schemas visible in src/main/butler/tools.ts and src/main/tools/docker-tools.ts. However, critical gaps in descriptions, parameter documentation, and error handling significantly impact quality. Butler tools (1-5) share identical 200+ character descriptions that are dense, specification-heavy prose aimed at internal state management rather than LLM decision-making. Docker tools (6-10) have minimal descriptions (14-50 chars), falling well below the 50-200 char baseline for LLM-optimized guidance. Parameter descriptions are present but vary widely in actionability. The schema structures are well-formed with proper types and enums, but descriptions lack practical guidance on when to call tools, what they return in detail, and how to recover from failure. No tool includes error classifications or recovery guidance patterns.
Read a file from the Docker mount
Create a default mode task for Butler. Use valid JSON only. taskKey must be unique within this turn. If user explicitly specifies a workspace directory, set workspacePath accordingly. When threadStrategy is "reuse_last_thread", optionally provide reuseThreadId for precise thread reuse. For follow-up reuse tasks, include this block inside initialPrompt: [Reuse Followup Message] ... [/Reuse Followup Message]. workspacePath supports absolute path or path relative to butler.rootPath. dependsOn references other taskKey values in the same turn. Treat the original user request as locked task body; do not rewrite it in initialPrompt. initialPrompt is execution addendum only: constraints, method, risk controls, or user habit notes. initialPrompt must not change scope/time/object/format/acceptance from the locked task body.
Create a email mode task for Butler. Use valid JSON only. taskKey must be unique within this turn. If user explicitly specifies a workspace directory, set workspacePath accordingly. When threadStrategy is "reuse_last_thread", optionally provide reuseThreadId for precise thread reuse. For follow-up reuse tasks, include this block inside initialPrompt: [Reuse Followup Message] ... [/Reuse Followup Message]. workspacePath supports absolute path or path relative to butler.rootPath. dependsOn references other taskKey values in the same turn. Treat the original user request as locked task body; do not rewrite it in initialPrompt. initialPrompt is execution addendum only: constraints, method, risk controls, or user habit notes. initialPrompt must not change scope/time/object/format/acceptance from the locked task body.
Butler task tool descriptions (tools 1-5) are dense technical specifications (200+ chars) rather than LLM-optimized guidance. They read like API documentation (spec versions, field constraints, internal reuse protocols) instead of explaining WHEN the LLM should call each tool or WHAT makes one mode (default vs ralph vs email) appropriate for a given request. The description for create_default_task occupies 450+ chars of parameter constraints, leaving no room for decision-making context. Expected baseline: 50-200 chars with clear selection guidance.
Docker operation tool descriptions (tools 6-10) are critically short (14-50 chars) and generic. 'Execute a shell command inside the Docker container' leaves an LLM unable to distinguish execute_bash from dozens of shell-execution tools. Missing: what types of tasks is this optimized for? What timeouts apply? What encoding does output use? When should the LLM use cwd vs env? Baseline requires 50-200 chars minimum with differentiation context.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 52 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Create a expert mode task for Butler. Use valid JSON only. taskKey must be unique within this turn. If user explicitly specifies a workspace directory, set workspacePath accordingly. When threadStrategy is "reuse_last_thread", optionally provide reuseThreadId for precise thread reuse. For follow-up reuse tasks, include this block inside initialPrompt: [Reuse Followup Message] ... [/Reuse Followup Message]. workspacePath supports absolute path or path relative to butler.rootPath. dependsOn references other taskKey values in the same turn. Treat the original user request as locked task body; do not rewrite it in initialPrompt. initialPrompt is execution addendum only: constraints, method, risk controls, or user habit notes. initialPrompt must not change scope/time/object/format/acceptance from the locked task body.
Create a loop mode task for Butler. Use valid JSON only. taskKey must be unique within this turn. If user explicitly specifies a workspace directory, set workspacePath accordingly. When threadStrategy is "reuse_last_thread", optionally provide reuseThreadId for precise thread reuse. For follow-up reuse tasks, include this block inside initialPrompt: [Reuse Followup Message] ... [/Reuse Followup Message]. workspacePath supports absolute path or path relative to butler.rootPath. dependsOn references other taskKey values in the same turn. Treat the original user request as locked task body; do not rewrite it in initialPrompt. initialPrompt is execution addendum only: constraints, method, risk controls, or user habit notes. initialPrompt must not change scope/time/object/format/acceptance from the locked task body.
Create a ralph mode task for Butler. Use valid JSON only. taskKey must be unique within this turn. If user explicitly specifies a workspace directory, set workspacePath accordingly. When threadStrategy is "reuse_last_thread", optionally provide reuseThreadId for precise thread reuse. For follow-up reuse tasks, include this block inside initialPrompt: [Reuse Followup Message] ... [/Reuse Followup Message]. workspacePath supports absolute path or path relative to butler.rootPath. dependsOn references other taskKey values in the same turn. Treat the original user request as locked task body; do not rewrite it in initialPrompt. initialPrompt is execution addendum only: constraints, method, risk controls, or user habit notes. initialPrompt must not change scope/time/object/format/acceptance from the locked task body.
Download a file from the Docker mount
Edit a file inside the Docker mount by replacing a string
Execute a shell command inside the Docker container
Upload a file into the Docker mount
No error handling or recovery guidance in any tool. Error responses will be raw exceptions without classification (retryable vs fatal) or actionable next steps. Example: if execute_bash times out, the agent has no guidance on whether to retry, increase timeout, or decompose the command. Example: if edit_file fails to find 'old_str', the agent gets no hint to call cat_file first to inspect the current content. Missing pattern: pattern:recovery-guide.
Output schemas are not documented for any tool. LLMs cannot plan downstream calls or extract structured data when tool responses are not described. Example: what does execute_bash return? stdout+stderr? exit code? Is the output a string or an object with fields? What does create_default_task return, a taskId? confirmation? Baseline requirement: 100% of A+ tools document return types.
Parameter descriptions are present but often lack actionable constraints. Example: 'cwd' parameter in execute_bash has description 'Working directory inside the container (optional)' but does not state: relative to what? does it affect path resolution? what is the default? 'encoding' in upload_file does not specify when to use utf-8 vs base64 or consequences of choosing wrong. Baseline: parameter descriptions should include format/range/valid values and decision guidance.
create_default_task and create_ralph_task accept mutually exclusive or interdependent parameters (e.g., reuseThreadId only valid when threadStrategy='reuse_last_thread') but do not declare this in descriptions. LLMs will pass both threadStrategy options and attempt invalid combinations. Baseline: document parameter relationships explicitly.
Butler task tools declare complex nested objects (handoff, expertConfig, loopConfig) as parameters but do not document the expected structure, required fields, or valid sub-field values. Example: loopConfig.trigger is typed as 'object' with 'description' as 'Schedule, API, or file-based trigger' but no structure definition. LLMs cannot construct valid payloads without explicit schemas for nested objects.
No tool declares what file paths, workspace paths, or resource identifiers are valid. Example: does workspacePath accept only absolute paths? does it support ~/ expansion? what are path traversal guards? cat_file and edit_file accept arbitrary paths but do not specify security boundaries (e.g., confined to container mount). Missing pattern: pattern:tool-gateway.
execute_bash and edit_file risk unintended data loss but lack confirmation patterns. A malformed regex in edit_file 'old_str' could match the entire file and replace it with 'new_str'. execute_bash can run destructive commands (rm -rf). No tool offers dry-run, confirmation, or reversibility. Missing pattern: pattern:confirmation-request.
download_file and upload_file accept 'encoding' as optional enum (utf-8 | base64) but do not specify what happens when encoding is omitted. Does the tool auto-detect? default to utf-8? fail? Baseline: defaults must not cause data loss; if omitting encoding is ambiguous, make it required or clearly document the default behavior.