MCP server for entertainment content recommendations providing tools for discovering and searching films and television shows, managing genres, and comparing LLM responses
Greenroom has 8 tools with mostly complete schemas and reasonable descriptions, but several critical gaps prevent a higher score. All tools start with action verbs (compare, list, discover, search) which is good. Most have documented input schemas with types and descriptions. However, descriptions vary widely in quality, some are specific and actionable (e.g., discover_films with detailed filter explanations), while others lack sufficient context for LLM selection (e.g., compare_llm_responses' description doesn't explain WHEN to use this tool vs alternatives). Output schemas are not formally documented in the source, they're inferred from TypedDict hints in models/responses.py. Error handling is minimal; there's basic validation in compare_llms but no recovery guidance or actionable error messages. The tool set is well-composed (each does one thing) and naming is clear, but parameter documentation could be more complete with explicit enums for sort_by fields and tighter constraints on numeric ranges.
Categorize all available genres by mood/tone. Groups entertainment genres into mood categories (Dark, Light, Serious, Fun) using a hybrid approach: hardcoded mappings for common genres with LLM-based categorization for edge cases and unknown genres.
Compare how multiple agents respond to the same prompt. As of 2026, this defaults to comparing a resampling of claude with a freshly generated response from Ollama.
Retrieve a list of films based on optional filters like genre, release year, original language, and sorting preferences. For now, defaults to TMDB service.
Retrieve a list of television shows based on optional filters like genre, first air year, original language, and sorting preferences. For now, defaults to TMDB service.
List all available entertainment genres across media types and providers.
Output schemas are not formally documented; inferred from function return type annotations only. LLMs cannot see what fields to expect from tool results at runtime, limiting their ability to plan follow-up calls and chain tools effectively.
Enumerated parameters use string descriptions instead of formal JSON Schema enums. sort_by parameters in discover_films/discover_television accept a fixed set of values ('popularity.desc', 'popularity.asc', 'vote_average.desc', 'vote_average.asc', 'date.desc', 'date.asc') but are typed as nullable strings with no enum constraint, allowing LLMs to hallucinate invalid sort options.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 67 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 54 | - | v1 |
Get a simplified list of available genre names. Uses LLM sampling to extract genre names from the full genre data, returning a clean, formatted list without IDs or media type flags. Falls back to direct extraction if sampling is not supported.
Looks up films by title. Use this when the user provides the title of a specific film. Use discover_films instead when browsing by criteria like genre or year. For now, defaults to TMDB service.
Looks up television shows by title. Use this when the user provides the title of a specific show. Use discover_television instead when browsing by criteria like genre or year. For now, defaults to TMDB service.
list_genres_simplified and categorize_genres reference sampling (LLM resampling, LLM-based categorization) but the server declares sampling=false. This creates a contract violation, the tool descriptions promise LLM-powered behavior that may not be available, causing silent fallbacks that could confuse agents about result quality.
compare_llm_responses description does not explain WHEN or WHY an LLM would call this tool. It states what it does but provides no decision context, an agent must guess whether to use this for testing, verification, or model comparison.
No structured error responses. While compare_llm_responses includes input validation with actionable messages, other tools lack documented error scenarios. Tools that call external APIs (TMDB) should document rate limits, authentication failures, and network timeouts with recovery guidance.
Output limits not documented. discover_films and discover_television accept max_results parameters but the tool descriptions don't state absolute caps. Without explicit limits (e.g., 'hard cap at 100 results per page'), large result sets could exceed token budgets and degrade LLM reasoning.
Parameter format constraints missing. display_language and original_language parameters accept ISO 639-1 codes but lack explicit format documentation (length, case, allowed values). Without this, LLMs may pass invalid locale strings.