This server has two tools with schemas and descriptions present, but both suffer from significant quality issues. Tool names use domain-specific prefixes (sandbox_initialize, sandbox_exec) but lack clear action verbs that follow the verb_noun pattern. Descriptions are lengthy and partly in Chinese, making them less effective for LLM reasoning. Parameter descriptions exist but contain example values (e.g., 'python:3.9-slim') which LLMs tend to reuse literally. Output schemas are not documented, responses are inferred from code inspection (dicts with 'success', 'error', 'container_id' fields) but the tool descriptions do not specify what fields to expect. Error handling returns basic dict responses with 'success' and 'error' keys, but provides no actionable recovery guidance. Security is a major concern: sandbox_exec accepts arbitrary shell commands, creating an injection vector. No input validation, sanitization, or permission gates are evident. The tools lack idempotency guarantees and confirmation patterns despite being destructive operations. Parameter naming is inconsistent: 'image' vs 'container_id' vs 'workspace_id' have no clear ID-type suffixes. No enum constraints for image selection.
在沙箱环境中执行命令。 在指定的容器中运行一个或多个shell命令并返回输出。 参数: container_id: 从initialize调用返回的容器ID commands: 要在沙箱环境中运行的命令列表 workspace_id: 工作区ID(用于持久化容器模式) 返回: 包含每个命令执行结果的字典
初始化一个新的代码执行计算环境。 基于指定的Docker镜像创建一个容器,默认使用Python slim镜像。 参数: image: 作为基础环境的Docker镜像 (例如 'python:3.9-slim') use_persistent: 是否使用持久化容器(如果为True,image参数将被忽略) 返回: 包含容器ID和工作区ID的字典,可用于与该环境交互
Tool names lack clear action verbs following verb_noun pattern. 'sandbox_initialize' and 'sandbox_exec' use domain prefix but do not start with precise action verbs (e.g., 'create_sandbox', 'run_command_in_sandbox'). This weakens LLM intent inference.
Descriptions contain example values ('python:3.9-slim', container ID format) which LLMs reuse literally rather than adapting to context. Rubric mandates: do not put example values in descriptions; use enum constraints or format declarations instead.
Output schema not documented in tool descriptions. Responses return dict with fields like 'success', 'error', 'container_id', 'workspace_id', 'workspace_path', 'mode', but LLM has no formal specification of these fields. Rubric: 100% of A+ tools have documented return types.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 35 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 0 | - | v1 |
sandbox_exec accepts arbitrary shell commands as a list of strings with no input validation, sanitization, or constraint. This is a command injection vector. Rubric: treat all agent-provided input as untrusted; sanitize against command injection; apply constraint validation early.
Error responses lack recovery guidance. A response like {'success': False, 'error': 'Docker service unavailable'} tells the LLM nothing about what to do next. Rubric: error responses must tell the LLM what to do next with actionable guidance.
Destructive operations (sandbox_exec, sandbox_initialize with container creation) lack confirmation pattern or dry-run support. Agents make mistakes, irreversible operations should support confirm_before_execute. Rubric: irreversible operations should support a dry-run or confirmation step.
No permission gates or scope declaration. Tools can execute arbitrary commands in Docker and create containers without checking caller authority. Rubric: gate destructive/sensitive tools behind permission checks; each tool should declare required permissions.
Descriptions are functional but suboptimal for LLM reasoning.
Parameter naming inconsistency: 'image', 'container_id', 'commands', 'workspace_id'. No clear ID-type suffixes applied uniformly. Rubric: when a parameter could be an ID, name, email, or position, suffix it with the type to avoid ambiguity.
'image' parameter has a default ('python:3.9-slim') but no enum constraint for valid image choices. Free-form string invites hallucinated or invalid Docker image names. Rubric: declare enum for known set of values; free-form strings invite invalid input.
No audit logging of tool calls. Code has logging statements for Docker operations but does not log who called the tool, with which parameters, at what time, and what happened. Rubric: agent-initiated actions must be traceable for compliance, debugging, and incident response.