Container Manager - manage Docker, Docker Swarm, and Podman containers. MCP+A2A Servers Out of the Box!
Container Manager MCP shows mixed quality across 7 tools. Naming is reasonably clear with action verbs (deploy_, stop_, get_, list_, cm_). Descriptions are present but often generic and lack specificity about when to use each tool vs alternatives. Parameter descriptions vary, some tools like deploy_specialist_container have decent detail, while others like cm_info_operations are sparse. Schema definitions are present in JSON but lack type information for complex fields (ports, env, labels, health_check are all strings requiring JSON parsing). No documented output schemas, critical omission for tools that return container objects, status dicts, or lists. Error handling is not visible in the provided code; no evidence of recovery guidance or error categorization. Tool composition splits concerns reasonably (deploy vs stop vs status) but the 'cm_*' tools (cm_info_operations, cm_image_operations, cm_list_hosts) feel overloaded with action parameters, cm_info_operations accepts 'get_version' or 'get_info', cm_image_operations accepts 4 different actions. This multi-action pattern makes tool selection ambiguous for LLMs. Security: no evidence of credential handling, permission checks, or audit logging in the visible code.
Manage container images: list, pull, remove, or prune images.
Manage container manager info operations.
List the host aliases you can pass as ``host`` to any cm_* operation to manage a REMOTE machine's Docker (via Docker-over-SSH). Every cm_* tool accepts ``host``: omit it to use the local Docker socket, or pass an alias here to target another machine — e.g. a swarm MANAGER node for cluster ops — without deploying a container-manager on each box. Aliases come from the tunnel-manager inventory (``~/.config/agent-utilities/inventory.yaml``) and connect as ``ssh://<user>@<hostname>:<port>``.
Deploys a specialist agent as a container using the Agent OS ContainerConfig schema. Pulls the image, creates the container with port mappings, env vars, labels, and optional health check, then starts it.
Gets the status and health of a specialist container.
Lists all containers managed by Agent OS (filtered by managed-by=agent-os label).
No documented output schemas for any tool. LLMs cannot predict what fields will be present in responses, forcing them to guess downstream field names and causing failed compositions.
Multi-action tools (cm_info_operations, cm_image_operations) use a single tool with an 'action' enum rather than separate tools per action. This makes tool selection ambiguous for LLMs (which tool should I call for pulling an image?) and violates single-responsibility principle.
Complex parameters (ports, env, labels, health_check) are JSON strings with no schema validation or examples. deploy_specialist_container expects ports as '{"8004": "8004"}' but LLMs frequently hallucinate wrong JSON syntax. These should be structured objects with explicit schema constraints.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 53 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Stops a running specialist container, optionally removing it.
Parameter descriptions are sparse. Most parameters lack guidance on allowed values, format constraints, or examples. E.g., 'health_check' accepts a shell command but there's no description of expected format, syntax, or timeout behavior.
No error handling or recovery guidance documented. If a container deploy fails, pull fails, or image not found, what should the LLM do next? No actionable error messages visible in code.
Tool names like 'cm_info_operations' and 'cm_image_operations' are non-standard and don't follow verb_noun convention. This makes intent ambiguous and difficult for LLMs to parse.
list_specialist_containers has no pagination parameters (limit, offset, next_cursor). If 100+ containers exist, LLM cannot paginate and may hit result size limits or token explosion.
manager_type parameters default to null with 'auto-detect' behavior. This is underspecified, what happens if auto-detect fails? Should the LLM provide an explicit value? No guidance.
No visible security patterns: no evidence of permission checks, audit logging, scope declarations, or input sanitization for injection attacks.