SkyPilot Code Sandbox MCP Server - A remote code execution service that provides sandboxed code execution across multiple programming languages with session management and resource pooling.
The server defines 6 tools with explicit FastMCP registration and clear action-verb naming. All tools have descriptions and documented parameters, placing this above the median. However, several critical gaps prevent a higher score: (1) parameter descriptions lack constraint information (ranges, enums, format requirements), (2) output schemas are not documented, tools return JSON strings but LLMs cannot see the structure, (3) error handling is minimal with no recovery guidance, (4) some descriptions are generic and could better explain WHY to use a tool vs alternatives. The execute_code tool is the strongest (good usage guidance), but others (get_health_status, get_pool_stats) are minimal. Average param description length is ~40-60 chars, below the 72-char baseline for A+ tools.
Close an execution session.
Create a new execution session.
Execute code in a sandboxed environment. IMPORTANT: To install Python packages, use the 'libraries' parameter instead of running 'pip install' commands in your code. This is more reliable and efficient. IMPORTANT: Always reuse the same session_id if it exists for consecutive code executions unless you need to change the language or libraries. This maintains state between executions (variables, imports, etc.) and improves performance.
Check the health status of the code execution service.
Get session pool statistics.
Get the list of supported programming languages.
Output schemas not documented. All tools return 'str' (JSON string) but LLMs cannot parse the structure. No field-level documentation for result shapes.
Parameter descriptions lack constraints. 'language' and 'session_id' have no validation hints. 'timeout' has no range (1-3600? 1-60?). 'libraries' format is vague ('array|string|null' in schema but description does not explain JSON string fallback).
Error handling lacks recovery guidance. Errors return raw strings like 'error: Connection refused' with no suggestion for what the LLM should do next (retry? choose different session? check service health?).
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 64 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 9 | - | v1 |
Minimal tool descriptions for read-only tools. 'get_health_status' is described as 'Check the health status of the code execution service.' (60 chars), too terse. Should explain: what does 'healthy' mean? When should an LLM call this? What are the failure modes?
'language' parameter has no enum constraint in schema (inferred as open string with default 'python'). LLM could pass invalid language like 'python3' or 'bash' without validation. Rubric requires enums for bounded choice sets.
Parameter relationship not documented: 'session_id' in execute_code should be tied to session lifecycle (create_session → execute_code → close_session). No cross-tool guidance for agents.
No default for 'timeout' parameter in execute_code is confusing. Schema says 'default: 30' but parameter docs say 'Execution timeout in seconds' with no range hint. Is 30 reasonable for all code? What's the max?