Multi-model MCP server for the multi-model research workflow. Single tool: query_llm_models(prompt, models, system_prompt?) for openai, google, and/or anthropic.
Single tool 'query_llm_models' has a moderately detailed description and visible input schema, but lacks formal type constraints, comprehensive parameter documentation, and output schema specification. The tool description is adequate (183 chars) but does not clearly state WHEN to use it vs alternatives, what it returns structurally, or dependencies. Parameters lack granular descriptions and enums. No error handling guidance is visible. This is a typical community-grade tool, functional but not production-grade.
Call one or more of OpenAI (GPT), Google (Gemini), and Anthropic (Claude) with the same prompt. models: any non-empty list of "openai", "google", "anthropic" (e.g. ["anthropic"], ["openai", "google"], or all three). Returns dict of model name to reply text.
Output schema not documented. Tool returns dict[str, str] but description does not explain the structure (keys are normalized model names; values are response text). LLM cannot plan downstream use without seeing the output shape.
Parameter 'models' lacks enum constraint. Description says 'openai', 'google', 'anthropic' are valid, but no formal enum is declared. LLM may hallucinate invalid values like 'gpt-4', 'claude', 'gemini' which the tool then normalizes. Should declare enum: ['openai', 'google', 'anthropic'].
Parameter descriptions are minimal or absent. 'prompt' lacks guidance on length/format. 'models' describes content but not the constraint that it must be non-empty (though code validates this). 'system_prompt' marked optional with no description of its effect or typical use cases.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 51 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 53 | - | v1 |
No error handling guidance visible in tool description. When API keys are missing (raised as ValueError), the error message is user-facing but does not guide the LLM on recovery. Should document: 'Requires OPENAI_API_KEY, GOOGLE_API_KEY (or GEMINI_API_KEY), and ANTHROPIC_API_KEY in environment.'
Tool description does not state that it modifies no state (safe to retry) or when to call it instead of calling a single model directly. 'Use this to compare responses from multiple providers' or 'Use when you need diverse perspectives or want to validate an answer' would clarify intent.