MCP server for executing coding tasks in sandboxes with OpenCode integration
sandbox-mcp demonstrates solid definition quality with well-structured Zod schemas, clear parameter documentation, and proper HTTP transport. All three tools have explicit definitions with input schemas and descriptions visible in src/agent/tools.ts. Naming follows verb_noun patterns (opencode_run_task, opencode_get_result, opencode_list_runs). Descriptions are adequate (40-80 chars each). However, output schemas are NOT documented in the source code, responses are formatted via formatToolResponse() but the actual return structure (fields, types) is not declared anywhere visible. Error handling is minimal: the codebase shows formatErrorResponse() exists but no guidance on error recovery or classification. Parameter constraints are well-defined (enums for status, min/max for limits), but descriptions lack dependency hints and context about when to call each tool vs alternatives.
Get the result of a completed run
List runs with optional filtering by session, status, and time
Execute a coding task in a sandbox. Creates session if needed, or continues existing session.
No documented output schemas. Tools use formatToolResponse() to return content, but the actual response structure (fields, types, nested objects) is not declared anywhere. LLMs cannot plan downstream tool calls or extract the right data without knowing what fields are present.
Minimal error handling and recovery guidance. formatErrorResponse() exists in tools.ts but no tool description mentions what errors can occur, how to recover, or whether failures are retryable. E.g., opencode_get_result doesn't document 'run not found' vs 'run still running' vs 'run failed' scenarios.
Ambiguous timestamp format in opencode_list_runs 'before' parameter. Description says 'Unix timestamp cursor' but does not clarify seconds vs milliseconds. This is a common LLM error, conversion mistakes cause off-by-factor-of-1000 bugs.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 62 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 45 | 1.25.1+ | v1 |
Terse descriptions for opencode_get_result and opencode_list_runs. At 27 and 59 chars respectively, they lack context on intent, when to call them vs alternatives, and what to expect. Baseline for A+ tools is 50-200 chars with explicit WHAT/WHEN/HOW structure.
No idempotency documentation. opencode_run_task creates sessions and runs tasks, are these idempotent if called twice with the same sessionId and task? Agents will retry on ambiguous failures; non-idempotent operations risk duplicate execution.
No tool annotation hints (readOnlyHint, destructiveHint, idempotentHint). opencode_get_result and opencode_list_runs are read-only and should be marked as such so agents prioritize them. opencode_run_task is write-heavy and should carry a destructiveHint or idempotentHint.