A Docker MCP Server for managing Docker containers, images, networks, and volumes
This Docker MCP server defines 22 tools with consistent naming (all action verbs), reasonable descriptions, and complete JSON schemas. However, several critical gaps prevent a higher score: (1) Output schemas are completely undocumented, the code uses docker_to_dict() but never shows what fields are returned, forcing LLMs to guess at response structure; (2) Most parameter descriptions are minimal (10-30 chars) and lack depth about constraints, valid ranges, or expected formats; (3) No error handling guidance, tools will fail silently or with raw Docker errors rather than actionable recovery instructions; (4) High-risk destructive tools (remove_container, remove_image, recreate_container) lack confirmation patterns or dry-run capabilities; (5) The execute_command tool accepts arbitrary commands with a 'privileged' flag but provides no input validation, injection protection guidance, or security warnings. Positive aspects: tool names follow verb_noun convention consistently (list_*, get_*, create_*, remove_*, etc.); input parameters use proper JSON Schema with types; common parameters are annotated with helpful type hints (ContainerID, ImageName, etc.); the Pydantic models (ListContainersFilters, etc.) show schema-aware design. The server is functionally complete for basic Docker operations but lacks the defensive documentation and error handling expected of production-grade agent tools.
Create a new Docker container
Create a new Docker network
Create a new Docker volume
Execute a command inside a running Docker container
Get details of a specific Docker container
Get details of a specific Docker image
Get details of a specific Docker network
Get details of a specific Docker volume
Output schemas completely undocumented. Code uses docker_to_dict() but no schema is documented for any tool's return value. LLMs cannot plan downstream calls or extract required fields (container_id, image_id, etc.). This violates pattern:tool and forces guesswork on response structure.
Destructive tools lack confirmation/dry-run capability. remove_container, remove_image, remove_network, remove_volume, and recreate_container are marked DESTRUCTIVE or IRREVERSIBLE but provide no confirm_before_execute pattern, dry-run flag, or safety gates. Agents can trigger data loss without review.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 65 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 35 | - | v1 |
List all Docker containers
List all Docker images
List all Docker networks
List all Docker volumes
Pull a Docker image from a registry
Recreate a Docker container (destroy and recreate with new configuration)
Remove a Docker container
Remove a Docker image
Remove a Docker network
Remove a Docker volume
Restart a Docker container
Run a Docker container (create and start)
Start a stopped Docker container
Stop a running Docker container
No error handling guidance. Tools will fail with raw Docker SDK exceptions or cryptic error messages (e.g., 'No such image'), but descriptions never explain what to do on failure. Pattern 'recovery-guide' requires actionable error responses like 'Image not found. Try pull_image() first to fetch it from a registry.'
execute_command lacks input validation and security warnings. Tool accepts arbitrary command strings and a 'privileged' flag with only 10-char descriptions. No mention of command injection risk, escaping requirements, or LLM guidance on safe vs unsafe patterns. Should warn about shell metacharacters and recommend explicit parameter arrays for multi-arg commands.
Parameter descriptions are minimal (10-30 chars) and lack actionable constraints. For example, 'command' param is described as 'Command to run in container' without mentioning format (shell string vs array?), escaping, or length limits. 'timeout' is described as 'Timeout in seconds' without min/max bounds. Pattern 'constrained-input' requires explicit ranges and formats.
Filter parameters use generic object types without nested schema documentation. list_containers and list_images accept 'filters' as an object with 'label' arrays, but the schema does not document how labels should be formatted ('key', 'key=value'?), whether multiple labels are AND or OR, or pagination support. LLMs cannot construct correct filter queries.
Result limits not documented. Tools like list_containers and list_images can return hundreds of items, but descriptions never state a limit or pagination strategy. According to pattern 'paginated-result', large result sets must be capped (recommend 20-50) and offer pagination or next_cursor. Without this, context window can be exhausted.
Missing chaining IDs in response documentation. create_container likely returns a container_id, but if subsequent tools (start_container, execute_command, remove_container) require container_id, responses must be documented to include it. Pattern 'response-shaper' requires downstream tool compatibility. Undocumented IDs force wasted lookup calls.
ports and volumes parameters accept complex nested objects without detailed format documentation. ports maps container ports to host ports, but description does not clarify format ('8080:80'? dict with 'ContainerPort' and 'HostPort'?). volumes can be list or dict, ambiguous. LLMs will guess and fail.
No indication of idempotency. Pattern 'idempotent-operation' requires tools to state whether repeated calls with the same parameters are safe. For example, create_container with the same name will fail on retry, should this tool be idempotent (use existing if present) or fail? Not documented.