MCP server for managing Docker Compose stacks via Dockge remote server
Dockge MCP presents a moderately well-structured server with 14 tools covering Docker Compose stack management. Strengths: all tools have descriptions (10 - 15 sentences typical length), input parameters are explicitly typed with JSON Schema, and tool registration is clear via fastmcp decorators. Weaknesses: descriptions are verbose but lack the terse, LLM-optimized 50 - 200 char target range; parameter descriptions are often generic ('The endpoint specifies which remote server...'); no output schemas are documented for the LLM; error handling relies on a generic 'ok' boolean flag without recovery guidance; several tools share nearly identical parameter sets (start_stack, stop_stack, restart_stack, update_stack) which creates cognitive load for LLM selection; no enum constraints exist even where they would be valuable (e.g., shell parameter in start_remote_terminal defaults to 'bash' but accepts any string). The terminal/logging abstraction (terminal_name return pattern) is reasonable for async operations, but the dependency on get_terminal_logs for all result inspection is not documented in output schemas. Parameter naming is inconsistent (stackName vs name; stack_name vs stackName across tools). Security concerns: endpoint parameter is treated as user input without validation, no evidence of auth/authorization checks, endpoint sanitization, or rate limiting documented in the visible code.
Stops the Docker Compose stack, remove all orphaned containers, deletes the compose and evn yaml files in the remote server.
Deploys a Docker Compose stack in a remote server.
Retrieves a list of all Docker networks from a remote server.
Retrieves the configuration, compose and env yaml of a single Docker Compose stack from a remote server.
Fetches the real-time status of services within a Docker Compose stack from a remote server.
Retrieves the last X lines from a terminal session (created by deploy_stack, delete_stack, start_stack, stop_stack, restart_stack, update_stack, or start_remote_terminal).
Parameter naming inconsistency across similar tools. Tools use 'stackName' (start_stack, stop_stack, restart_stack, update_stack, get_stack_service_status) vs 'name' (deploy_stack, save_stack, delete_stack) vs 'stack_name' (start_remote_terminal). LLMs struggle with these inconsistencies when composing multi-tool plans.
No output schemas documented. Tools return dicts with generic fields like 'ok', 'msg', 'terminal_name', but LLMs cannot plan downstream operations without knowing the result structure. E.g., deploy_stack returns {ok, msg, terminal_name}, will get_terminal_logs always have data immediately, or must the agent wait? What fields does get_stack return? What do list_stacks and get_docker_network_list return beyond a generic dict?
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 52 | - | v1 |
Retrieves the a list of Docker Compose stacks from a remote server.
Restart a Docker Compose stack on a remote server.
Saves the Docker Compose stack files (`compose.yaml` and `.env`) in a remote server without deploying it.
Sends a command to a remote interactive terminal session started by 'start_remote_terminal'. Use bashful commands like 'ls', 'cd', 'cat', etc. to interact with the service's container. Use heredoc syntax (<<EOF ... EOF), sed, or echo with redirection (>>) to create or modify files. Long running commands may not return output until they complete. Send SIGINT (Ctrl+C) to stop a running command. Repeat if you don't see log updates. After sending input, the updated terminal screen can be fetched using 'get_terminal_logs'.
Starts an interactive terminal session for a service in a Docker Compose stack running in a remote server.
Start a stopped Docker Compose stack on a remote server.
Stops a running Docker Compose stack on a remote server.
Runs docker compose full to update images for given stack running in a remote server and redeploys it.
list_stacks and get_docker_network_list have no input parameters documented. No pagination (limit, offset, page), no filtering, no result count declaration. If these return hundreds of stacks/networks, context window explodes. Rubric requires pagination support and result limits stated in description.
Descriptions are verbose (15+ sentences) but lack LLM-optimized precision. 'Deploys a Docker Compose stack in a remote server.' could be 'Deploys a Docker Compose stack to a remote Dockge server. Pass compose.yaml content, .env variables, stack name, and target endpoint. Returns terminal_name to track deployment logs via get_terminal_logs.' Rubric baseline shows A+ tools average 194 chars; these are 2 - 3× that.
start_remote_terminal has 'shell' parameter defaulting to 'bash' as a free-form string with no enum. Shell choices should be restricted (bash, sh, zsh, ash) to prevent invalid values. Rubric: 'When a parameter accepts one of a known set of values, declare it as an enum.'
Error handling is minimal. Tools return {ok: bool, msg: str} but provide no error classification (retryable vs user-fixable vs fatal), no recovery guidance, and no categorization of failures. E.g., if delete_stack fails with 'Stack not found', the LLM learns nothing about whether it should retry, search for the stack, or report to the user.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) despite clear risk classification visible in the assessment. Tools marked DESTRUCTIVE (delete_stack) and WRITE (deploy_stack, save_stack, etc.) should carry explicit annotations to guide agent planning and safety checks.
endpoint parameter is required or defaults to empty string for all tools but is never validated in visible code. No evidence of auth checking, rate limiting, or sanitization against command injection/path traversal. Agents pass untrusted user-provided endpoints directly to socket.io calls without validation.
Destructive operations (delete_stack) have no confirmation step or dry-run option. Agents can inadvertently delete production stacks. Rubric: 'Irreversible operations should support a dry-run or confirmation step.'
Several tools (deploy_stack, save_stack, etc.) accept large string inputs (composeYAML, composeENV) with no documented size limits, format validation, or injection checks. LLMs could pass malicious or oversized payloads.