A dynamic MCP server that automatically loads and registers Python tools from a tools directory, specifically designed for Docker management operations
This server has 12 Docker management tools with visible schemas and descriptions, but quality is inconsistent. Most tools have basic descriptions (in Chinese) and parameter schemas, but descriptions are often generic and lack LLM-optimized guidance on WHEN to use each tool and what makes it distinct from similar tools. Several tools exhibit overlapping functionality (docker_container_action does 4 things: start/stop/restart/remove) that should be split. No structured output schemas are documented, functions return free-form markdown strings rather than JSON objects with typed fields. Error handling is minimal (mostly 'try/except' returning error strings without recovery guidance). Security concerns: docker_exec_run and docker_pip_install allow arbitrary shell command execution with minimal input validation. No rate limiting, no audit logging, no permission gates. Chinese descriptions limit LLM clarity, descriptions should be in English or bilingual for multi-language support.
对容器执行操作:启动、停止、重启、删除。
从容器内部复制文件到 AstrBot 本地存储。用于提取代码沙箱生成的结果(图片、文档等)。
删除镜像
在运行中的容器内执行命令 (相当于 docker exec)。
获取容器日志 (后 N 行)
查看容器的详细信息(IP、挂载、环境变量等)。
列出所有容器。
docker_container_action violates single-responsibility principle, accepts 4 mutually exclusive actions (start/stop/restart/remove) as a string enum. This combines 4 distinct operations into one tool, forcing LLMs to reason about which action to invoke and making the tool harder to select correctly.
No output schemas documented. All tools return free-form markdown strings (e.g. '📦 **容器列表...'). LLMs cannot structure data extraction or plan downstream tool calls when output format is undefined. Responses should be JSON with typed fields (containers: [{id, name, status, image}], total: number).
Descriptions are in Chinese and often generic/minimal. E.g. docker_list_images: '列出本地镜像' (13 chars) provides no context on WHEN to call it vs docker_inspect_container, or what the output contains. Descriptions should be 50-200 chars, in English, explaining purpose, use case, and output structure.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 45 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 42 | - | v1 |
列出本地镜像
在指定容器内快速安装 Python 依赖库 (自动使用清华源加速)。
拉取/下载镜像
重置镜像(强制更新到最新版并清理旧缓存)。
运行一个新的容器 (docker run)。
docker_exec_run and docker_pip_install accept arbitrary shell commands with minimal input validation. No escaping, no sandboxing, no command whitelisting. An LLM could be tricked into passing 'rm -rf /' or exfiltrating secrets. Requires command validation, whitelisting, or confirmation before execution.
Error handling is minimal. Most errors return a raw string like '❌ 操作失败: {str(e)}'. LLMs receive no guidance on whether to retry, what went wrong, or how to fix it. Errors should be structured with severity (retryable/user_fixable/fatal) and actionable recovery hints.
No audit logging or permission gates. Tools can delete images, remove containers, and execute arbitrary code without logging who called them, when, or what happened. Destructive/sensitive operations require permission checks and audit trails for compliance.
docker_run_container accepts 'ports' as a free-form string that is parsed as Python literal. The description example '{\'80/tcp\': 8080}' with escaped quotes is confusing. Should accept structured input (JSON object or key-value pairs) validated server-side, or use a dedicated ports_mapping parameter with explicit schema.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint). LLMs cannot distinguish read-only vs write vs delete operations at a glance. docker_delete_image and docker_container_action (remove) should have destructiveHint=true. docker_list_* should have readOnlyHint=true.
docker_get_logs and docker_inspect_container return massive unstructured data. docker_inspect_container dumps raw JSON with all environment variables, mount points, etc. Should limit output to essential fields (id, status, IP, primary_image) and offer separate tools for detailed inspection.