MCP server for LLM/VLM model selection — compare 300+ models with real-time benchmarks, pricing, and personalized recommendations. No API key required.
The server defines 4 focused, read-only tools with clear naming and reasonable descriptions. All tools follow verb_noun patterns (model_info, list_top, compare_models, recommend). Descriptions are present and moderately detailed (100-200 chars each), explaining what data is returned and when to use each tool. However, input parameter descriptions are sparse or missing, only the top-level parameter has a description in most cases, lacking detail about format, constraints, and valid ranges. Output schemas are undocumented: the rubric requires documentation of return types (pattern:tool, Dimension 1.D), but no explicit schema definitions appear in the source. The tool definitions are visible and properly registered via registerXxxTool() functions, so inference penalty does not apply. Error handling is present (errorReporting=true) but not visible in the provided code excerpts, preventing full assessment of recovery guidance quality.
Compare 2-5 LLM/VLM models side-by-side: pricing, benchmarks, capabilities. Returns a compact Markdown comparison table (~400 tokens).
List top-performing models for a category: coding, math, vision, general, cost-effective, open-source, speed, context-window, or reasoning. Returns up to 10 ranked models with benchmarks and pricing.
Get detailed information about a specific LLM/VLM: pricing, benchmarks, capabilities, release date. Returns ~200 tokens.
Get personalized LLM/VLM recommendations based on your use case, requirements, and constraints. Considers performance, cost, latency, modalities, and more. Returns 3-5 ranked recommendations with reasoning.
Output schemas not documented. The rubric (pattern:tool, Dimension 1.D) requires tools to document their return types, but no explicit output schema definitions are visible in the source code. Each tool description mentions what kind of data is returned (e.g., 'Returns ~200 tokens', 'Returns up to 10 ranked models', 'Returns a compact Markdown comparison table'), but the actual JSON structure is undocumented. LLMs need to know what fields to expect in responses to plan downstream calls and extract data correctly.
Input parameter descriptions lack detail and constraints. Most parameters have only a bare description; none specify format, range, valid values, or examples for the constraint. For example, 'model' in model_info is described as 'Model ID or partial name (e.g., ...)' but does not state minimum/maximum length, allowed characters, or whether fuzzy matching is supported. 'category' in list_top lists valid enum values in the description but not as a formal constraint, inviting LLMs to hallucinate invalid categories.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 45 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 38 | - | v1 |
recommend tool has too many loosely-defined parameters. 'useCase', 'priority', and 'constraints' are all free-form strings. 'priority' should be an enum (speed|cost|quality|reasoning|vision|context as stated in description), and 'constraints' is vague, LLMs may struggle to format it correctly. The tool would benefit from structured inputs: priority as enum, constraints as sub-parameters (budget, latency_ms, min_context_tokens), and useCase as a typed field or enum of common scenarios.
No pagination or result limits explicitly declared in schema. list_top accepts a 'limit' parameter (default 10, max 10 in description) and compare_models accepts 2-5 models, but minItems/maxItems constraints appear only in the description, not in the JSON Schema itself. Formal schema constraints are machine-parseable and prevent LLMs from passing invalid values.
Error handling code not visible. errorReporting=true suggests error handling exists, but the provided source excerpts (index.ts, package.json, openrouter.ts fetcher, smoke test) do not show tool implementations or error response handling. Cannot verify that errors include recovery guidance (pattern:recovery-guide) or actionable context per the rubric (Dimension 1.E).