Host inspection and system monitoring MCP server for AIOps
The AIOps MCP server provides 7 well-intentioned system monitoring tools with clear naming conventions and reasonable descriptions. However, there are significant gaps in parameter documentation, output schema documentation, and error handling guidance. Most tools follow verb_noun naming (get_*, check_*, list_, analyze_), which is good. Descriptions exist for all tools and most are between 100-200 chars, meeting baseline expectations. However, parameter descriptions are inconsistent, some parameters lack descriptions entirely (e.g., analyze_log_file's 'log_path' has a description but 'max_lines' and 'error_keywords' lack detailed constraints). Output schemas are inferred from return type hints in docstrings but are not formally documented in the tool definitions visible in the source. Error handling is minimal, most tools return structured responses but lack guidance on what the LLM should do when a tool fails (e.g., when a log file doesn't exist or a network connection times out). No tool declares permissions, implements idempotency hints, or includes dependency information. The server provides good coverage of system monitoring patterns but lacks the rigor expected for production use.
Analyze a log file for errors and return relevant information.
Check network connectivity by attempting to connect to a specific host.
Check if a specific port is open on the given host.
Check if a specific process is running.
Get basic system information.
Get basic system metrics including CPU, memory, and disk usage.
List all running services/processes with their resource usage.
Parameter descriptions incomplete or missing. analyze_log_file's 'max_lines' and 'error_keywords' parameters lack detailed constraint documentation (min/max for integers, format examples for enums). check_network_connectivity's 'timeout' parameter lacks range constraints.
Output schemas not formally documented in tool definitions. Descriptions include return types as prose (e.g., 'Dict containing...') but the actual JSON Schema structure is not visible in the @mcp.tool() decorator or input/output schema blocks. LLMs cannot parse prose schema descriptions reliably.
Error handling lacks recovery guidance. Tools return error states (e.g., 'connected: False', 'open: False') but do not guide the LLM on what to do next. analyze_log_file may fail silently if a file doesn't exist, no guidance on checking file path or trying an alternative log location.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | B | 73 | 2026-07-28+ | v2 |
| 2026-03-09 | C | 65 | - | v1 |
No tool annotations (readOnlyHint, idempotentHint). All 7 tools are read-only operations but do not declare this via tool annotations. Without explicit annotations, agents cannot confidently retry or optimize these calls.
list_running_services hard-caps results at 20 and sorts by CPU, but the limit is not documented in the tool description. LLMs may assume they're receiving all processes and make incorrect decisions. The description should state: 'Returns top 20 processes by CPU usage; use other monitoring tools to see complete process list.'
get_system_info description is vague ('Get basic system information'). No detail on what fields are returned (OS, hostname, uptime, Python version, etc.). Compare to get_system_metrics, which clearly lists CPU, memory, disk.
No permission declarations. Tools read system state (processes, ports, logs, network) but do not declare required permissions (read:process, read:network, read:filesystem). Agents cannot be configured with least-privilege scope.