Turnkey immutable and isolated Linux sandboxes for MCP servers. Provides multiple pre-configured MCP servers (sqlite, shell, filesystem, fetch) running in hardened Docker containers with strict isolation, read-only root filesystems, dropped capabilities, and optional network access.
mcp-box provides 5 tools across shell execution and SQLite database operations. While tool names follow verb_noun conventions and descriptions exist, the implementation has critical gaps: (1) Parameters lack type descriptions, input schemas show only parameter names and basic types with minimal guidance on format/constraints; (2) No output schemas are documented, responses are returned as plain text strings rather than structured objects, making it impossible for LLMs to reliably extract and chain data; (3) Error handling is generic, 'Error executing command' tells agents nothing about retry strategy or recovery paths; (4) No parameter validation descriptions, e.g., run_command accepts any bash command with no guidance on safety, timeouts, or expected output format; (5) Output is unstructured text, list_tables and describe_table return formatted text tables rather than JSON objects with typed fields, violating the structured-output pattern. The tools are functional but lack production-grade definition quality.
Describe the schema (columns, types, nullability, defaults) for a specific table.
List all tables available in the database schema.
Execute a read-only SELECT query on the SQLite database and return results as formatted text.
Execute a bash shell command inside the isolated container sandbox and return stdout/stderr.
Execute a modifying query (INSERT, UPDATE, DELETE, CREATE, DROP) on the SQLite database.
No output schemas documented for any tool. All responses are unstructured plain text (formatted tables or concatenated strings). LLMs cannot reliably extract structured data for chaining or validation.
Parameter descriptions are minimal or missing. 'command: string', 'sql: string', 'table_name: string' lack guidance on format, constraints, length limits, or safety boundaries. LLMs have no actionable constraints.
Error handling is generic and non-actionable. Errors like 'Error executing command: [exception]' or 'Error: Command execution timed out' do not guide recovery. No classification (retryable vs. fatal) or next-step guidance.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | F | 48 | 2026-07-28+ | v2 |
write_query is destructive but has no dry-run, confirmation, or transaction rollback pattern. An agent can accidentally DELETE all data with no safety guard.
No pagination support documented for list_tables or read_query. If a database has thousands of tables or a query returns millions of rows, responses could exhaust context windows. No limit or offset parameters.
run_command accepts arbitrary bash commands with no input validation or injection prevention documented. LLMs could be tricked into passing dangerous payloads (e.g., $(rm -rf /), $(exfil_data)). No sanitization guidance.
Tool descriptions lack context on when to use each tool vs. alternatives. No hint that read_query is for SELECT and write_query is for INSERT/UPDATE/DELETE; LLMs must infer from names.