Async Ansible utilities with inventory and playbook helpers
This Ansible MCP server has moderate structural quality but significant gaps in descriptions, parameter documentation, and error handling. All 8 tools are explicitly defined with REST endpoints and Pydantic input models, providing basic schema structure. However, tool descriptions are minimal (mostly 5-15 words), parameter descriptions lack context, and error responses are generic. The tool composition follows verb-noun naming (list_*, run_*, generate_*, validate_*) which is appropriate, but output schemas are undocumented, callers cannot plan downstream operations. Security concerns exist: no input sanitization visible for file paths (path traversal risk in generate_playbook, validate_playbook), no permission gates on destructive operations (run_ad_hoc, run_playbook, generate_playbook), and timeout handling is incomplete. Error messages are raw stdout/stderr without actionable recovery guidance. The server scores above F-range due to explicit tool registration and basic parameter typing, but falls short of C-range (60+) due to missing descriptions, undocumented schemas, and weak error handling.
Generate and write Ansible playbook file to disk
Get Ansible version
List all hosts from inventory
List Ansible inventory structure and hosts
Perform ping test on all hosts in inventory
Execute Ansible ad-hoc command on specified host
Execute Ansible playbook with Server-Sent Events streaming
Tool descriptions are extremely minimal (5-15 words). Rubric baseline for production tools is 194 chars average; these descriptions range 20-60 chars and lack WHEN to call them, WHAT they return, and HOW they differ from similar tools. Examples: 'List Ansible inventory structure and hosts' (41 chars) does not explain whether it returns a parsed JSON structure or raw text, or when to prefer list_inventory vs list_hosts. LLMs cannot reliably select the right tool.
Parameter descriptions are missing or trivial. Rubric requires EVERY input parameter to have a non-empty description explaining what it controls. Examples: (1) run_ad_hoc.module, described only as 'Ansible module name' (no examples, no link to Ansible module docs, no guidance on common modules like 'command' vs 'shell'); (2) run_ad_hoc.args, described as 'Module arguments' (what format? YAML? key=value? quoted?); (3) run_playbook.extra_vars described as 'Extra variables to pass to playbook' (should clarify: JSON object syntax, escaping rules, size limits). These gaps force LLMs to guess.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 45 | - | v1 |
Validate Ansible playbook syntax
Output schemas are completely undocumented. Rubric requires 'Document the output schema. LLMs need to know what fields to expect so they can plan downstream tool calls and extract the right data.' All 8 tools return dict/JSON (via Pydantic .dict() or asyncio subprocess output), but the source code includes zero docstrings or inline comments describing the structure. Examples: (1) list_inventory returns either an error dict or parsed JSON, what fields are in the inventory structure? (2) run_playbook returns Server-Sent Events, what format? (3) run_ad_hoc returns {stdout, stderr, returncode}, should return stdout be piped into a follow-up call? No guidance.
Path traversal vulnerability in generate_playbook and validate_playbook. Code attempts sanitization via os.path.basename() but STILL exposes risk. Example: generate_playbook accepts file_name input, calls os.path.basename(input.file_name.strip()), then writes to PLAYBOOKS_DIR/sanitized_name. An attacker could pass '../../../etc/passwd', basename would strip leading ../ but if PLAYBOOKS_DIR is a relative path, the write could still escape. Rubric: 'Treat all agent-provided input as untrusted. Sanitize against SQL injection, command injection, and path traversal.' validate_playbook is similar, it reads a file path from input.playbook with minimal validation. Agents, if tricked by prompt injection, could read arbitrary files.
No permission gating on destructive operations. Rubric: 'Gate destructive or sensitive tools behind permission checks. Verify the calling user/agent has authority before executing.' Tools run_ad_hoc, run_playbook, and generate_playbook perform WRITE/DESTRUCTIVE actions but include zero permission checks. Code shows no authentication or authorization, any caller can invoke run_playbook to execute arbitrary Ansible playbooks on production infrastructure. No scopes declared, no audit trail, no user/agent identity checks.
Error responses lack actionable recovery guidance. Rubric: 'Error responses must tell the LLM what to do next: ... A raw error code or stack trace gives the agent nothing to act on.' All error responses in the code are raw: {stderr: stderr.decode(), returncode: proc.returncode}. Examples: (1) validate_playbook returns proc.returncode 1 with stderr text, LLM cannot determine if syntax is fixable or if inventory file is missing; (2) ping_hosts timeout returns generic 'Command timed out after 30s', no guidance to retry, skip hosts, or reduce inventory; (3) list_inventory 'Inventory file not found: X', no suggestion to check paths or available inventories. No error classification (retryable vs user-fixable vs fatal).
No input validation before subprocess execution. Rubric: 'Do not assume LLMs will follow constraints. Validate inputs early and return clear error messages.' Examples: (1) run_ad_hoc accepts host, module, and args as free-form strings and passes them directly to subprocess, no validation that host exists in inventory, module is valid Ansible module, or args is safe YAML/JSON syntax; (2) run_playbook.extra_vars is typed as Optional[Dict[str, Any]], no size limit, no value type checking, could contain malicious payloads; (3) run_playbook and run_ad_hoc do not check if ansible/ansible-playbook binaries exist or are accessible. Validation happens only at execution, with poor error messaging.
Missing pagination/limits on list tools. Rubric: 'Tools returning lists should accept page/offset and limit parameters and return a total count or next_cursor. Without pagination, large results blow the context window.' list_inventory, list_hosts do not cap output. In large Ansible inventories (hundreds of hosts, thousands of group memberships), the JSON response could easily exceed 100KB, killing context windows. No limit parameter, no pagination, no next_cursor. Baseline for production tools: 20-50 item limit with pagination support.
Inconsistent error response format across tools. Some return {error: 'message'}, others {stderr: '', returncode: 0}, others raw subprocess output. Rubric: 'Return structured objects with typed fields.' LLMs need a predictable error format. Example: list_inventory returns {error: 'file not found'} but validate_playbook returns {stdout: '', stderr: 'error', returncode: 1}. This inconsistency forces LLMs to parse multiple formats and risks missing error conditions.
No audit logging or action traceability. Rubric: 'Log who called what tool, with which parameters, at what time, and what happened. Agent-initiated actions must be traceable for compliance, debugging, and incident response.' Code includes zero audit logging. When run_playbook executes, there is no record of: who triggered it, when, what inventory/playbook were used, what changed on target hosts. Critical for infrastructure automation, required for compliance and incident response.