docker-mcp has 4 tools with explicit definitions visible in server.py. Tool naming follows verb_noun conventions (create-container, deploy-compose, get-logs, list-containers), which is good. However, descriptions are severely underdeveloped, all 4 tools have descriptions under 100 characters, ranging from 36 - 60 chars, which violates the rubric baseline of 194 chars average and fails to provide actionable context for LLM tool selection. Parameters lack per-parameter descriptions entirely, the inputSchema JSON includes property descriptions in the schema itself (e.g., 'Docker image name', 'Container name'), but these are sparse and not comprehensive. Schema quality varies: create-container and deploy-compose have well-formed schemas with required fields; get-logs and list-containers are minimal. Error handling is present but generic, the catch block returns a flat TextContent with 'Error: {str(e)}', offering no recovery guidance per the pattern:recovery-guide. Output schemas are not documented, the code returns List[types.TextContent] but no structured response format is specified, violating the pattern:response-shaper requirement. No pagination is implemented for list-containers, even though it's a discovery tool that could return large result sets. No input validation beyond argument presence checks. No tool-specific error categorization (retryable vs. fatal). Security: the server does not validate or sanitize Docker Compose YAML input before passing it to the executor, risking injection attacks. Composition: tools are single-purpose and properly separated (list, get, create, deploy), which is correct. The prompts feature is implemented, which is a positive feature for guiding LLM behavior, but that doesn't offset core definition quality gaps.
Create a new standalone Docker container
Deploy a Docker Compose stack
Retrieve the latest logs for a specified Docker container
List all Docker containers
Tool descriptions are critically short (36 - 60 chars) and lack actionable context. Rubric baseline is 194 chars; descriptions must explain WHAT the tool does, WHEN to use it, and WHAT it returns. Current descriptions are too terse for LLM tool selection.
Input parameters lack per-parameter descriptions in the tool definition. While the JSON schema includes brief property descriptions (e.g., 'Docker image name'), they are minimal and do not guide LLM reasoning. For example, 'ports' parameter lacks explanation of format (host:container mapping), and 'environment' lacks guidance on variable naming conventions.
Output schema is not documented. Tools return List[types.TextContent] with no specification of response structure, field names, or what data consumers should expect. Violates pattern:response-shaper and makes it difficult for agents to extract and chain results.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 51 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 43 | - | v1 |
Error handling is generic and non-actionable. The catch block returns 'Error: {str(e)} | Arguments: {arguments}', which gives the LLM no recovery guidance and may expose internal stack traces or sensitive arguments. Does not follow pattern:recovery-guide or pattern:error-classification.
deploy-compose tool accepts raw Docker Compose YAML as a string parameter without validation or sanitization. YAML injection attacks are possible. The tool should validate YAML syntax and structure before executing, and should reject YAML that mounts sensitive host paths or runs privileged containers without explicit user confirmation.
list-containers tool lacks pagination support. If many containers are running, returning all results in one response violates mxe:enforce-result-limits (cap at 20 - 50 items) and risks exhausting context windows. Tool should accept limit and offset/cursor parameters and return a total_count.
No input validation beyond argument presence checks. For example, create-container does not validate that 'image' is a valid Docker image reference, that port mappings follow host:container format, or that environment variable names are alphanumeric. Invalid input produces cryptic Docker API errors instead of clear guidance.
Irreversible operations (create-container, deploy-compose) have no confirmation or dry-run mode. A misbehaving agent could launch many containers or stacks unintentionally. Should implement a confirm-before-execute pattern or at least a dry_run parameter.