MCP server for the 1001 Albums Generator challenge. Provides tools to explore your listening history, analyse taste from ratings and reviews, compare group members, and get context on today's assigned album.
This server demonstrates strong structural foundations with 18 well-defined tools, comprehensive descriptions, and proper schema usage via Zod. However, there are significant gaps in output schema documentation and error handling guidance. All tools are read-only (low risk), which simplifies error recovery but the server lacks explicit guidance on what to do when tools fail. Tool naming is consistently verb-first and clear. Descriptions are detailed and actionable (averaging 180-220 characters), explaining WHEN to use tools and what they return. Parameter descriptions are present and mostly specific. Critical gap: output schemas are not documented in the source code, while Zod definitions exist internally, the MCP tool registration does not expose return type information that LLMs can reference. This forces LLMs to infer output structure from context. Additionally, error handling returns generic error text without recovery guidance (e.g., 'try search_users() first'). The server would improve significantly with explicit return type schemas and error recovery hints.
Deep taste comparison between two users in a group. Returns shared albums, mean absolute divergence in ratings, a similarity score (0–100), and whether their taste is compatible.
Retrieve a project's full album history with optional filtering, sorting, and pagination. Results include album metadata, your rating, review text, and global community ratings.
Check the health and status of the caching layer (Redis or in-memory).
Get all-time community favourites and least favourites: highest and lowest rated albums, sorted by average rating or controversy. Use to see where you diverge most from the pack.
Retrieve global community statistics: the highest and lowest rated albums across all users, sorted by average rating, controversy, or vote count. Use this to contextualize a user's taste — e.g. "you rated the community's #1 album a 2, which is bold."
Find the albums that divided your group most (highest standard deviation in ratings) and the ones you all agreed on (lowest standard deviation). Helps spark discussions or highlight surprising consensus.
Output schemas not documented in MCP tool registration. While Zod schemas exist internally in the implementation, the server does not expose return type information to the MCP protocol layer. LLMs cannot see what fields to expect from tool responses and must infer structure from context.
Error responses lack recovery guidance. In register-tool.ts, errors are caught and returned as generic text ('Error: {message}') without telling the LLM what to do next or whether the error is retryable. For example, if an API returns 404 for a project, the agent should know to call a search tool first.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | <=2025-11-25 | v2 |
| 2026-03-09 | C | 65 | - | v1 |
Discover which genres you rate highest on average, and how many albums of each genre you've encountered. Use this to understand your taste breadth and depth — e.g. "you're a rock devotee with 300+ albums rated" or "you have a surprising love of electronic music despite low exposure".
Retrieve all group members' ratings and reviews for a specific album.
Compute pairwise similarity across all group members. Returns a compatibility matrix showing who has the most and least similar taste, and average similarity for each member.
Retrieve group metadata: name, members, current/latest album, all-time highest and lowest rated albums.
Convenience wrapper: get the latest album assigned to a group and all members' reviews for it.
Visualize your entire listening journey as a narrative arc — how your taste evolved, where you took big stylistic leaps, and when you hit peak enthusiasm or critical periods. Returns segments, milestones, and trend data suitable for charting.
Quick stats on a project: number of albums generated, rated, and unrated.
See how your ratings are distributed across the 1–5 scale. This reveals whether you're a harsh rater, generous, or somewhere in between — and how consistent you are.
Search, filter, and analyse reviews you've written. Find reviews matching a query (album name, artist, genre, text content), optionally filter to a specific album, and get back a curated set with metadata like your rating vs. community rating.
Discover which styles (more granular than genres) you rate highest on average, and how many albums of each style you've encountered.
Analyses a project's full listening history and returns a structured taste profile for the user. Use this as a starting point for any conversation about a user's music taste, listening patterns, or identity as a listener. It is intentionally broad — use it to orient yourself before reaching for more specific tools. The profile includes: - DECADE DISTRIBUTION: How many albums per decade, and which decade dominates. Use this to open discussions like "you're clearly a 70s rock person" or "your history skews surprisingly modern". - TOP GENRES & STYLES: The genres and styles that appear most frequently across the history. Note that genres are broad (e.g. "Rock") while styles are more specific (e.g. "Psychedelic Rock") — both are useful at different levels of conversation.
Retrieve today's assigned album for a user or group, including full metadata, release date, genres, and links to streaming services.
Missing error classification. The server does not categorize errors as retryable, user-fixable (e.g., invalid input), or fatal. Without this, agents cannot reason about failure recovery, whether to retry, ask the user, or escalate.
Parameter 'limit' lacks explicit bounds. Several tools (get_review_insights, get_community_stats, get_community_ratings) accept a 'limit' parameter with no documented minimum or maximum. This allows LLMs to request absurd values (e.g., limit=999999) that could cause timeout or memory exhaustion.
get_cache_status description is too brief (48 characters). 'Check the health and status of the caching layer (Redis or in-memory).' does not explain WHEN to call this tool or what the LLM should do with the result. Should expand to ~150+ characters with context about operational use.
Pagination consistency unclear. get_album_history and get_community_stats both support 'offset' and 'limit' parameters, but it is not documented whether omitting 'limit' returns all results or a default page size. Agents need explicit defaults to avoid fetching unbounded lists.
projectIdentifier parameter documentation incomplete. Several tools accept 'projectIdentifier' (project name or sharerId), but the schema does not specify the format, length, or example. Is it case-sensitive? Alphanumeric only? This ambiguity invites invalid input from LLMs.