Execute commands interactively on remote Windows machines using the WinRM protocol
evil-winrm-py provides a functional Windows remote execution toolkit with 4 tools that have explicit schemas, descriptions, and input validation. Tool naming follows verb-noun conventions (winrm_login, list_sessions, winrm_execute, winrm_logout), which is correct. However, descriptions are brief (averaging ~90 chars, below the 194 baseline), parameter documentation is minimal, and output schemas are not explicitly documented. The server implements tool annotations (destructiveHint for winrm_execute), which is modern. Error handling exists but lacks actionable recovery guidance. Parameter defaults are well-chosen (port=5985, ssl=False, uri='wsman') and do not cause data loss. The critical risk is that session management relies on integer session_ids without clear documentation of how sessions are created, stored, or cleaned up, this could lead to agent confusion or stale sessions.
List all active WinRM sessions and their session_id.
Run a command on a WinRM-connected Windows host.
Authenticate to a remote Windows host over WinRM. Returns the session_id to pass to winrm_execute and winrm_logout.
Close an active WinRM session.
Output schemas not documented. LLMs cannot predict what list_sessions, winrm_execute, and winrm_logout return, forcing them to guess at response structure and blocking downstream tool composition.
Descriptions are too brief (avg ~90 chars, well below 194 baseline). 'List all active WinRM sessions and their session_id' lacks context on when to call it or what the session_id is used for. 'Run a command on a WinRM-connected Windows host' omits whether the command is PowerShell, batch, or another dialect, and does not explain what happens on error.
winrm_execute and winrm_logout both have optional session_id parameters with the note 'Optional if only one session is active.' This creates ambiguity: what happens if zero or multiple sessions exist and session_id is omitted? Error message? Default to first session? This invites silent failures.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | B | 71 | 2026-07-28+ | v2 |
winrm_execute lacks any indication of what success looks like or how errors are surfaced. Does it return exit code, stdout, stderr separately? Is there a max output size? No documentation, forcing LLMs to guess and potentially misinterpret response structure.
Authentication parameters (priv_key_pem, cert_pem) accept file paths as strings rather than using secure credential injection. If an agent logs parameter values, certificate paths (and potentially sensitive key material) could leak into logs. Consider environment variable injection or vault integration.
No error handling guidance. If winrm_login fails (bad credentials, network timeout, certificate error), the description does not suggest what the agent should do next or how to disambiguate the failure cause.
winrm_execute is marked destructiveHint=true, which is correct, but the description does not emphasize this or offer a dry-run pattern. Agents may not fully appreciate the irreversibility of command execution without explicit warning in the description text.