Read-only Model Context Protocol server wrapping HomeLab Monitor's HTTP endpoints for remote exploration of homelab infrastructure: hosts, containers, systemd services, GPU, AI models, alerts, and costs.
HomeLab Monitor MCP server demonstrates strong definition quality overall. All 19 tools are explicitly registered with FastMCP, have clear verb-noun naming conventions (list_, get_, scan_), and include substantive descriptions (average ~180 chars). Input schemas are present for all tools with proper type definitions and parameter descriptions. However, there are gaps in output schema documentation, while responses are mentioned in descriptions, formal JSON Schema output specs are not visible in the source code. The server is read-only (all tools are READ_ONLY risk), which simplifies security concerns but limits composability. Tool composition is well-designed: list_hosts → get_host → specialized tools (get_snapshot, get_containers, etc.) follow a clear discovery pattern. Parameter naming is consistent (range, path, status, run_id). Some tools have implicit dependencies (e.g., get_entity_cost expects 'kind'/'name' pairs from get_costs output) that could be more explicit. Error handling is mentioned but not visible in the source excerpt.
AI model servers: which models are loaded, their VRAM use, and who is driving them (caller→server connection-seconds attribution over `range`, e.g. "6h", "24h", "7d"). Answers "why is the GPU pinned, and which service is calling it?".
Current / recent alerts: which thresholds are breached now (GPU VRAM > X%, host RAM/CPU > Y%, disk < Z GB free) and which were recently (last 24h). Answers "what's broken right now, and what was broken?".
Full detail for one benchmark run by `run_id` (from `get_benchmarks`): model, context lengths tested, per-context {throughput_tok_sec, prefill/decode latency, fit/spill breakdown by layer}, and the GPU it ran on. An unknown id returns an HTTP 404 error.
Stored LLM benchmarks (Benchmark Lab tab): per model per context length, the measured tokens/sec throughput, what fits in VRAM vs spills to RAM, and optimal context cap. Filter by `range` and `model` (partial match ok). Answers "what can this GPU do, and how fast?".
Full Docker container list (Containers tab): per container name, image, state, health, exposed ports, RAM (`mem_bytes`), VRAM (`vram_bytes`), image disk and uptime — plus the summary counts.
Output schemas not documented in visible source code. While descriptions mention what each tool returns (e.g., 'returns total/free bytes and a nested entries tree'), formal JSON Schema output specifications are not visible. LLMs cannot reliably extract multi-level nested structures without explicit documentation.
No visible error handling or recovery guidance in source excerpt. Descriptions state 'An unknown id returns an HTTP 404 error' (get_experiment, get_benchmark) but do not guide the LLM on what to do next or how to recover.
Implicit parameter dependencies not documented. get_entity_cost expects 'kind'/'name' pairs from get_costs breakdown, but this dependency is not stated in the parameter descriptions. Forces LLM to infer structure.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | B | 74 | 2026-07-28+ | v2 |
Power-cost summary (Costs tab): what the machine drew and what it cost over `range`, with the live `tariff`, the `machine` totals (now/energy/cost for today, 7d, 30d and the range) and a ranked `breakdown` of which processes, containers, services and models cost the most. Pass `host` (a registered host's name from list_hosts) to price a remote machine from its own poll history — remotes have no per-process `breakdown` yet. Answers "what did my homelab cost, and what's the biggest line item?".
Cost drill-down for one process/container/service/model by `name` (use a `kind`/`name` pair from `get_costs`' breakdown; `kind` is optional). Returns its `energy_kwh`, `cost`, `avg_w`/`peak_w` and resource use over `range`. Answers "what did *this* model or container cost me?".
Alert/event log: OOM kills, threshold crossings (GPU VRAM pressure, host disk near-full), service crashes, etc., with the container/process/service that triggered it and when. Answers "what went wrong, and when?".
Full detail for one tracked run by `run_id` (from `get_experiments`): its logged-metric series (the loss curve, and a tokens/sec curve too if logged), the GPU power/util time-series over the run, and the priced energy it burned. An unknown id returns an HTTP 404 error.
Tracked training/eval runs (Experiments tab), each priced with the real GPU energy it burned. Optionally filter by `status` (running/finished/failed/killed). Each row carries its params, latest metrics (loss/accuracy…, plus throughput as `tokens_per_sec` when an LLM training script logged it), duration and cost. Answers "which runs ran, how did they do, what did each one cost, and how many tokens/sec were they pushing?".
GPU detail: current utilisation, VRAM used/total, power and temperature, plus per-model VRAM use and the caller→server attribution over `range`. Answers "why is the GPU pinned, and which service is calling the model server?".
Charted time-series the dashboard graphs over `range`: timestamps + aligned arrays for GPU util/VRAM/power/temp and host CPU/RAM/load/temp. Use for trends.
Full System / Network / Security inventory for a single host. `name` is the host's registered name, or "local" for the hub itself.
Installed-models registry (#219): every AI model available on the hub — grouped by provider (ollama, vllm, llama.cpp, LM Studio, ComfyUI, …) — not just what's currently loaded. Ollama entries carry on-disk size/quant/param detail; other providers carry name/loaded/vram. Answers "what can I run, and where?" without SSHing in to run `ollama list` or poke a server's API by hand.
RAM breakdown behind the memory treemap: per-service peak/avg/present RAM over `range`, the current RAM per process right now, kernel (non-reclaimable) memory, and totals. Answers "what is eating RAM on this hub?".
Full systemd unit list (Services tab): per unit active/sub state, description, listening ports, RAM, uptime, admin/watched flags and status verdict — plus the summary counts.
Current live vitals across the hub: GPU, host RAM/CPU, Docker and systemd health summaries (with any problem containers / failed units), pending OS updates and whether a monitor update is available. DB-free and cheap.
List every host in the fleet with headline vitals (hub listed first). Returns the roster: name, online status, OS, CPU/RAM load, fullest disk and any OS-upgrade/reboot hint per host. Start here, then drill in with `get_host`.
WizTree-style nested folder-size treemap for a host path (Disks tab). Wraps the monitor's on-demand scan and polls until done. `path` is an absolute host path (e.g. "/", "/var"); set `rescan=True` to force a fresh scan. Returns total/free bytes and a nested `entries` tree of {name, path, bytes, children}.
get_installed_models and get_ai_models have overlapping semantics ('which models are loaded/available'). Descriptions do not clearly distinguish when to call each. Similar names risk LLM confusion.
No pagination or result limits documented for tools returning large collections. get_containers, get_services, get_experiments (with optional status filter) could return hundreds of items. Descriptions do not state whether results are paginated or capped.