A Docker-based polyglot code execution environment supporting Python, Rust, Go, and Bun. Provides tools for workspace initialization, container lifecycle management, file operations, and code execution with isolated sandboxed environments.
Sunaba provides 7 tools with complete input schemas and verb-based naming. All tools have descriptions. However, descriptions are often terse (10 - 50 chars), lacking WHEN/WHY guidance and context for LLM selection. Parameter descriptions are present but minimal. Output schemas are not documented in the code. Error handling is present but lacks recovery guidance. Naming is clear (workspace_init, env_start, env_stop, env_clean, file_write, exec_code) with good action verbs, but descriptions don't explain state transitions or prerequisites (e.g., what happens if you call env_start on an already-running container?). The server has proper input validation (enums for language, required fields), but responses and error messages are not visible in the provided code.
Remove container and optionally workspace
Ensure and start language container
Inspect container status
Stop running container
Execute code inside the language container
Atomically write a file in the workspace
Create base workspaces for languages
Tool descriptions are too brief (10 - 50 chars) and lack WHEN/WHY context. Examples: 'Create base workspaces for languages' (38 chars) tells WHAT but not WHEN to call it or what prerequisites exist. LLMs cannot reliably decide if a tool is appropriate without deeper guidance.
No output schemas documented in code. Responses for all tools are not specified, preventing LLMs from understanding what fields to expect (e.g., does env_status return container state, health, uptime?). Output documentation is required for tool composition and downstream chaining.
env_start and env_clean have destructive/reversible side effects but descriptions do not explicitly state this. 'Ensure and start language container' does not warn that this modifies system state. Agents need explicit state-modification declarations to reason about idempotency and retries.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 50 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Error handling strategy is not visible in provided code. No recovery guidance in error messages (e.g., if env_start fails, should the agent call workspace_init first? Inspect logs?). Pattern: recovery-guide suggests errors should include actionable next steps.
Parameter descriptions are minimal and do not explain constraints or formats. Example: 'rel_path' in file_write says 'Relative path within the language workspace' but does not specify: what characters are allowed? Are symlinks resolved? What is the max path length? Does it support nested creation? The description should mention create_parents=true enables nested directory creation.
env_clean has a destructive operation (remove_workspace=true deletes files) but no confirmation or dry-run support. If an agent misspells the language or calls this accidentally, data is lost. Pattern: confirmation-request suggests destructive operations should support confirmation or dry-run modes.
exec_code has many optional parameters with implicit defaults. 'entrypoint defaults to language-specific default (e.g., main.py for Python)', but what is the default for Rust, Go, and Bun? These defaults are not specified in descriptions, forcing the LLM to guess or make unsafe assumptions.