A Model Context Protocol for managing and interacting with multiple virtual machines over SSH
ssh-mcp-py has 4 tools with complete schemas and descriptions, but quality varies significantly. All tools have proper input schemas with types, and most descriptions are adequate (ranges 60-180 chars). However, parameter descriptions are sparse or generic, error handling lacks recovery guidance, and there is no output schema documentation. Tool naming is mostly clear but could be more precise (e.g., 'execute_ssh_command' vs. 'run_command'). The server follows basic tool definition patterns but misses several production-quality signals: no tool annotations, no pagination support, no explicit error categorization, and limited guidance on when to use which tool. No sensitive parameter injection issues detected (no secrets exposed). Overall falls into the 'Fair' category with noticeable gaps in parameter documentation and output structure.
Run a shell command on a remote host over SSH. Args: hostname: Host alias from SSH config. command: Shell command to run. timeout: Connection/command timeout, seconds (default 30, max 300). max_length: Max stdout/stderr length in the response (default 1000, max 10,000,000); full output is in the log file. Returns: Command output (markdown), with exit code and the log file path.
Get config details (hostname, port, user, key) for a host alias. Args: hostname: Host alias from SSH config. Returns: Host configuration details.
List all host aliases from the SSH config. Returns: Newline-separated list of configured hosts.
Test SSH connectivity to a host without running a command. Args: hostname: Host alias from SSH config. timeout: Connection timeout, seconds (default 30, max 300). Returns: Whether the connection succeeded or failed.
Missing output schema documentation. No tool describes what fields are returned or their types. LLMs cannot plan downstream operations without knowing response structure.
Parameter descriptions are generic or incomplete. For example, execute_ssh_command's 'timeout' parameter lacks guidance on what happens at timeout (retry? fail?), and 'max_length' doesn't explain the trade-off between response size and usability.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint). execute_ssh_command is clearly destructive (WRITE risk) but has no annotation to signal this to the agent.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 69 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 46 | - | v1 |
Error handling lacks recovery guidance. If SSH connection fails, what should the agent try next? Test connection first? List hosts? Descriptions should guide agent behavior on failure.
Tool composition could be clearer. The description for execute_ssh_command mentions 'full output is in the log file' but doesn't explain when or how to retrieve logs. Agents may not know if log retrieval is needed or how to access logs.