Real-time system monitoring dashboard API for infrastructure metrics collection, storage, and visualization with alerting capabilities
This HomeLab Infrastructure Monitor server exposes 18 tools over HTTP/FastAPI with generally well-structured definitions. Most tools have descriptions (15/18) and input schemas with type information. However, several critical gaps reduce overall quality: (1) Missing or weak descriptions on 3 tools; (2) Inconsistent schema completeness, some tools have enums and format constraints (list_alerts, create_alert_rule) while others lack them (websocket_metrics); (3) No documented output schemas despite returning structured data; (4) Parameters lack descriptions in several tools (e.g., websocket_metrics 'action' parameter has no description of what happens when you call subscribe vs unsubscribe); (5) Destructive operations (delete_alert_rule, delete_host, cleanup_old_metrics) lack confirmation patterns or dry-run guidance; (6) No error handling documentation or recovery guidance in any tool. The server falls into the 'Fair' (C) to 'Good' (B-) range, solid foundation but needs refinement for production LLM agent use.
Acknowledge an alert.
Clean up metrics older than specified days. Admin endpoint for maintenance.
Create a new alert rule.
Register a new host and generate API key.
Delete an alert rule.
Delete a host and all its metrics.
Get host details by ID.
No output/response schemas documented for any tool. LLMs cannot infer what fields to expect, preventing downstream tool chaining and forcing them to guess at return structure.
Destructive operations (delete_alert_rule, delete_host, cleanup_old_metrics) lack confirmation patterns, dry-run modes, or explicit recovery guidance. Agents could delete critical infrastructure without safeguards.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 64 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 47 | - | v1 |
Get the latest metrics for all hosts. Returns the most recent metric of each type for each host.
API health check endpoint.
Ingest metrics from an agent. Requires agent API key authentication.
List all alert rules.
List alerts with optional filters.
List all registered hosts.
Query historical metrics with filters. Supports filtering by host, metric type, and time range.
Mark an alert as resolved.
Update an alert rule.
Update host metadata.
WebSocket endpoint for real-time metrics streaming. Clients can subscribe to host-specific metric updates and receive alerts, status changes, and metric data in real-time.
websocket_metrics tool has weak parameter descriptions. The 'action' enum and 'host_id' parameter lack detail on what subscribe/unsubscribe/ping actually do, when host_id is required, and what the output format is.
No error handling guidance in any tool description. Tools do not explain what to do if a host is not found, a metric query returns no results, or a write operation fails. LLMs will not know how to recover.
Several tools have weak descriptions (under 65 chars or generic text): resolve_alert ('Mark an alert as resolved', no context on when/why), delete_alert_rule ('Delete an alert rule', no guidance), cleanup_old_metrics ('Clean up metrics', vague).
No rate limits or throttling hints documented. A runaway agent could flood the monitoring system with ingest_metrics or query_metrics calls, degrading service.
Pagination parameters (skip/limit) are present but max limits are not enforced in descriptions. query_metrics defaults to limit=1000, that could return 1000+ records, bloating context.
websocket_metrics is a streaming/stateful endpoint but tool definition treats it as a simple request/response. WebSocket semantics (persistent connections, subscriptions, push messages) are not modeled as MCP tools should be.