AI-powered MCP server for managing a self-hosted media stack (Jellyfin, Sonarr, Radarr, qBittorrent, PyLoad)
Mediabox MCP presents 9 tools with consistently well-structured schemas and descriptions. Most tools declare action enums clearly and include pagination parameters where appropriate. However, there are critical gaps in error handling guidance, missing output schema documentation, and some parameter relationships are underdocumented. Tool names are domain-specific (media_query, series, movies, downloads) but lack clear action verb prefixes that would aid LLM disambiguation. Several tools bundle multiple conceptual concerns (e.g., library_ops combines scan, create, move, delete, list, rename, refresh) rather than splitting into focused single-purpose tools. The confirm_token pattern for destructive operations (library_ops, maintenance) is a strong safety signal, but lacks explicit guidance on what success looks like. Parameter descriptions generally range from adequate to good (60-150 chars), exceeding the rubric minimum, but output schemas are entirely absent from the specifications provided.
Manage downloads. direct=download from direct HTTP URL or YouTube/video sites ONLY (never for Google Drive, Mega, MediaFire). add=send to PyLoad for file hosters (Google Drive, Mega, MediaFire, etc.). status=check PyLoad queue + download folders. organize=move completed to Jellyfin. delete_pyload=remove PyLoad packages. Sonarr/Radarr/qBit: list_queue, cancel, purge, clean_orphans.
Manage files and libraries. scan=refresh Jellyfin. create=new library. move=move files/folders. delete=cross-layer delete (Jellyfin+Sonarr/Radarr+disk). list=browse files. rename=standardize episode names. refresh=refresh metadata. action=delete is two-step: first call returns preview + confirmToken; pass that token back in your next call to apply.
Server maintenance and cleanup. action:'check_jobs'=view background job status. action:'cleanup'=remove stale files/cache/temp/metadata/orphans (two-step: preview + confirmToken).
Search or list Jellyfin content, or get details of a specific item. action:'search' to find/list media (omit query to list all, use page/pageSize to paginate). action:'details' for seasons/episodes of a series (use seasonNumber to filter one season, page/pageSize to paginate episodes).
Manage movies via Radarr. action:'search'=find by name (results include movieId only when the movie is already in Radarr — `inRadarr:true`). action:'add'=add to monitoring (needs addTmdbId from search). action:'status'=view movies/queue/history. action:'remove'=delete. action:'releases'=find torrents. action:'grab'=download a specific release. ID rule: tmdbId comes from search and is ONLY for action:'add' (as addTmdbId). For releases/grab/remove use `movieId` (the Radarr internal id from a prior search where inRadarr:true, or from the response of action:'add'). The router auto-resolves a tmdbId passed as movieId, but only if the movie has already been added — if not it fails with a clear hint.
Output schemas completely absent. Tools define input schemas with good detail, but zero documentation of what each tool returns, field types, or structure. LLMs cannot plan downstream calls or extract needed data without response schemas.
Tool names lack action verbs. 'media_query', 'series', 'movies', 'downloads', 'optimize', 'maintenance' are nouns or gerunds. LLMs cannot infer intent from names alone. Should be verb_noun: 'search_media', 'manage_series', 'manage_movies', 'download_media', 'optimize_media', 'check_maintenance'.
library_ops bundles 7 unrelated actions (scan, create, move, delete, list, rename, refresh) into one tool. Violates single-responsibility principle. Each action should be a separate tool: scan_library, create_library, move_files, delete_media, list_files, rename_episodes, refresh_metadata.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 69 | 2026-07-28+ | v2 |
Optimize media files and fix subtitles. action:'analyze'=preview encoding/audio/subtitle issues. action:'optimize'=transcode. action:'fix_subs'=harmonize subtitle format/encoding.
UI helper tool presented to the LLM. Exposes domain-specific actions available in the conversation context (what can be done next).
Manage TV series via Sonarr. action:'search'=find by name (results include sonarrId only when the show is already in Sonarr). action:'add'=add to monitoring (needs addTvdbId from search). action:'status'=view series/episodes/calendar/missing/queue/history. action:'remove'=delete. action:'releases'=find torrents. action:'grab'=download a specific release. ID rule: pass `seriesId` (the Sonarr internal id, or a tvdbId — both are auto-resolved) for releases/grab/remove/episodes-view; never pass tvdbId blindly to grab/releases without searching first.
Server status and activity log. action:'status' for full overview (disk, libraries, sessions, users). action:'activity' for recent playback history.
Missing error handling guidance. No documented error responses, recovery steps, or actionable error messages. E.g., if 'search_media returns 0 results, should LLM try a partial query? If add_series fails due to duplicate, what's the next step?
Parameter relationships underdocumented. series/movies tools accept different parameters based on action, but schema constraints don't enforce mutually exclusive groups. E.g., 'addTvdbId' only valid for action='add', but nothing prevents passing it for action='search'.
Enum parameters lack descriptions. 'monitor' in series tool has enum values (all, future, missing, none, firstSeason, lastSeason) but no explanation of what each means. Same for 'quality' enum across series/movies.
'present_choices' tool is poorly defined. Description is vague, return structure undefined, unclear if this is a legitimate tool or meta-artifact. If it's for UI scaffolding only, it should not be in the MCP tool list.
confirmToken pattern used but under-explained. library_ops and maintenance both use two-step delete/cleanup with confirmToken, but descriptions don't clarify: what is the token format? How long is it valid? What happens if you pass the wrong token?
Missing pagination limits. Some tools accept 'limit' or 'pageSize' without documented minimums/maximums. E.g., 'pageSize' in media_query defaults to 50 but is unbounded, LLM could pass 99999 and cause performance issues.