One MCP server for the whole self-hosted media stack: Radarr, Sonarr, Prowlarr, Bazarr, Jellyfin, Seerr, SABnzbd, Transmission.
arr-mcp demonstrates good tool design with clear, specific action verbs and well-structured schemas. All three tools follow verb_noun naming conventions (diagnose, get_indexers, profile_issues). Descriptions are detailed and context-aware, explaining WHEN to use tools and how they handle edge cases. Parameter types are properly declared with zod schemas. However, there are notable gaps in parameter descriptions, output schema documentation visibility, and error handling guidance that prevent a higher score.
Why is this not playable? Walks the whole chain — requested, managed, monitored, downloaded, indexed, imported, scanned — and names the first thing that explains the absence, with what to do about it. Give a title as `query` for how a person actually asks; give `service` plus `id` only when you already have an exact item in hand (e.g. from get_media_details) — the explicit id wins if both are given. Works with services down: any step it could not check sets `certain: false` and the summary says what was missed, rather than guessing across the hole.
Prowlarr indexer health: which indexers are enabled, which are temporarily disabled and why, per-indexer query and grab counts, and — at detail: full — the queries indexers recently rejected and the reasons they gave. Failure messages and rejection reasons come from the indexer itself and are fenced as untrusted data.
Quality profile diagnostics across configured Radarr, Sonarr, and Whisparr instances, plus Profilarr drift detection when configured.
Output schemas incompletely documented in visible source code. For 'diagnose', the outputSchema definition is truncated mid-line. For 'get_indexers' and 'profile_issues', no output schema visibility in provided snippets.
Pagination parameters (limit, offset) lack min/max constraints. No documented upper bounds prevent agents from requesting excessive result sets that waste tokens and degrade LLM reasoning. Best practice: limit max 50-100, offset >= 0.
Detail parameter in get_indexers and profile_issues uses enum-like description ('minimal, normal, full') but is NOT declared as JSON Schema enum constraint. LLMs cannot parse natural-language enums reliably, should use proper enum type in schema.
Inferred effective spec: 2026-07-28+.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | B | 71 | 2026-07-28+ | v2 |
Several parameters lack concrete format/constraint guidance: 'id' format in diagnose (UUID? integer? opaque string?), 'instance' and 'profile' matching rules in profile_issues (exact match? case-sensitive?), 'service' resolution logic in diagnose.
Error handling guidance missing. No descriptions of what errors the tools return, how to recover, or what the LLM should do on failure (retry? ask user? refer to logs?). diagnose mentions 'certain: false' for incomplete checks but other tools lack failure mode documentation.