MCP server that delegates design tasks to the top-ranked model on designarena.ai via OpenRouter
The server defines 4 tools with explicit names, descriptions, and JSON Schema input definitions. All tool names follow verb_noun convention (get_*, query_*), descriptions are present and reasonably detailed (100-300+ chars each), and all parameters have type constraints and descriptions. However, there are notable gaps: (1) enum descriptions are excessively long and repetitive (category parameter descriptions repeat the full enum list 4 times across 4 tools, totaling 1000+ chars per tool), (2) output schemas are not documented, the source code returns text-wrapped responses but does not specify what that text contains or its structure, (3) error handling is basic (generic try/catch), and (4) no tool annotations (readOnlyHint, destructiveHint) despite read-only classification in metadata. Tools are well-scoped and composable, but output documentation is a significant gap for agent planning.
Get the current #1 design model from designarena.ai's crowdsourced leaderboard, optionally filtered by category. Returns model info and OpenRouter availability.
Browse the full designarena.ai design model rankings with optional category filter and pagination. Shows Elo ratings, win rates, and OpenRouter availability.
Send a design prompt to the best available model on OpenRouter, automatically selected from designarena.ai rankings. Skips models not available on OpenRouter. Requires OPENROUTER_API_KEY.
Send a design prompt to a specific model via OpenRouter. Accepts either an OpenRouter model ID (e.g. 'anthropic/claude-sonnet-4-5-20250514') or a Design Arena model name (e.g. 'claude-sonnet-4-5'). Requires OPENROUTER_API_KEY.
Output schema not documented. All 4 tools return `{ content: [{ type: 'text', text }] }` but the structure and content of `text` is not specified. Agents cannot plan downstream uses or validate responses.
Enum parameter descriptions are excessively verbose and repetitive. The 'category' parameter in each tool repeats the full category list and description (~1000+ chars) rather than referencing a shared definition. This wastes tokens and makes the server definition difficult to maintain.
No tool annotations despite read-only classification in metadata. Tools should include readOnlyHint=true in their definitions to signal safety to agents and enable optimization.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-21 | D | 59 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 50 | - | v1 |
Error handling is minimal: generic try/catch with message interpolation. No recovery guidance, no categorization (retryable vs user-fixable), no suggestions for next steps. Errors like 'API returned 401' give agents no direction.
Missing parameter validation details. Parameters like 'max_tokens' and 'temperature' have range constraints in schema but descriptions could be more explicit about behavior at boundary values (e.g., 'At temperature=0, output is deterministic').