MCP server for system metrics monitoring including CPU, memory, disk, network, processes, Docker containers, and thermal data with background sampling and alerting capabilities
This is a solid community MCP server with 19 well-named tools covering system metrics. Naming is excellent (all verb-noun pattern, clear action verbs like get_*, start_*, stop_). Descriptions are present for all tools and range from 80-160 chars, appropriate length. Input schemas are properly defined with types and descriptions visible in the code. However, there are notable gaps: (1) No explicit output schema documentation visible in the source, the code returns data structures but no formal schema is declared in tool registration; (2) Error handling descriptions are sparse, tools don't explain what errors are possible or how to recover; (3) Some parameters lack validation guidance (e.g., 'devices' param in get_disk_io_metrics has no format hints); (4) No per-tool audit trail or permission checks visible. The monitoring tools (start_monitoring, stop_monitoring, get_alerts) are well-composed and offer recovery patterns. Overall, this server demonstrates good naming discipline and functional parameter design, landing in the upper-middle range.
Get threshold alerts generated by the monitor since the last read.
Get CPU usage, temperature, and load average
Get disk I/O statistics including read/write throughput, IOPS, and I/O time
Get disk usage statistics for mount points
Get per-device disk I/O throughput rates (bytes/sec and IOPS) since the last sample.
Get Docker container metrics including CPU, memory, network, and block I/O usage
Get memory usage statistics including RAM and swap
Output schemas not formally documented in tool registration. While handlers return structured data (evident from code like json.Marshal calls), no explicit schema declaration is visible in the mcp.NewTool() registrations. LLMs cannot reliably infer return structure without documented schemas.
Error handling lacks recovery guidance. Handlers do not document what errors can occur or how agents should respond. For example, get_docker_metrics may fail if Docker daemon is unreachable, but no error description guides the agent to retry, check Docker status, or fall back to get_system_health.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 60 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 32 | - | v1 |
Get recent metric snapshots captured by the monitor.
Get the current monitoring status, latest snapshot, and buffered alerts.
Get active network connections with local/remote addresses, status, and owning PID
Get network interface statistics
Get per-interface network throughput rates (bytes/sec) since the last sample.
Get list of running processes sorted by resource usage
Get systemd service status for specified services
Get an aggregated system health dashboard with CPU, memory, disk, and uptime in a single call
Get system information including hostname, OS, uptime, and platform details
Get thermal status including temperatures and throttling information
Start background sampling of system metrics. Returns a snapshot and enables history and alerts.
Stop background sampling of system metrics.
Parameter validation rules underspecified. Parameters like 'devices' (get_disk_io_metrics) and 'mount_points' (get_disk_metrics) state they accept comma-separated values but lack hints on format, case sensitivity, or example values. LLMs may pass malformed input.
Deprecated Sampling feature enabled. Code calls s.EnableSampling() in main.go. Per current MCP spec (2026-07-28), server-initiated sampling is deprecated. Migrate to stateless request handling and let clients pull alerts via get_alerts() on demand.
Tool composition: get_alerts() separates alerting from metric tools, but start_monitoring/get_alerts require a multi-step agent flow (start → get_metrics_history → get_alerts). Consider bundling a high-level 'describe_system_anomalies' tool that calls the monitor internally and returns alerts + metrics in one response, reducing round-trips.