Sandy provides 6 tools with mixed quality. Tool definitions are visible and mostly include schemas, but descriptions are inconsistent in depth and usefulness for LLM selection. Naming follows verb_noun convention (sandy_*, prime), which is good. However, schemas lack comprehensive parameter descriptions, and several tools have vague or missing guidance on error conditions and prerequisites. The server implements a specialized sandbox execution domain (TypeScript in Docker + AWS SDK access), which is niche and well-scoped, but tool composition and chaining could be clearer.
Return the MCP SKILL.md content
Run a health check (baseline or connect). Uses an ephemeral session.
Create a session and return its scripts path
Create or delete the Sandy sandbox image
Resume an existing session and return its scripts path
Run a TypeScript script in the Sandy sandbox
Incomplete parameter descriptions across tools; critical execution semantics undocumented. sandy_run lacks explanation of script path resolution, arg passing mechanism, and return format. sandy_check does not document the relationship between 'action' type (baseline vs. connect) and optional imdsPort parameter.
Vague and opaque tool name ('prime') does not convey action or purpose. Violates verb_noun naming convention, reducing LLM's ability to infer intent from the name alone. Should be renamed to 'get_sandbox_capabilities' or 'get_skill_definition'.
No documented error handling or recovery guidance. Tools do not explain failure modes (e.g., What if a session expires? What if the image build fails? What if the script times out?) or what the LLM should do next (retry, call discovery tool, report to user).
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | D | 59 | <=2025-11-25 | v2 |
Missing tool composition guidance. No documentation of expected workflow (e.g., 'Call sandy_create_session first, then sandy_run'). Tools like sandy_resume_session and sandy_create_session lack clarity on when to call which, and tool outputs do not document what fields downstream tools require.
Output schemas not documented. The rubric requires tools to document return types and field names so LLMs can plan downstream calls. sandy_create_session and sandy_resume_session return 'scripts path' but no JSON schema is provided showing the actual response structure (is it a string, an object with a 'path' field, etc.).
STDIO transport only. Server is not remotely accessible and cannot be used by hosted MCP clients.