MCP server for integrating multiple AI APIs (OpenAI, Anthropic, Google Gemini, xAI Grok)
This MCP server has 5 tools with reasonable naming (verb-noun convention) and descriptions present, but significant gaps in schema completeness, parameter documentation, and output specification. The `chat` tool has the most detailed schema definition visible in the code, with proper parameter types and descriptions (messages, model, provider, temperature, max_tokens, stream). However, the remaining 4 tools (list_models, compare, analyze, generate) show incomplete schema definitions, parameter types are not visible in the source code provided, only inferred from context. The server returns error responses as dictionaries without structured guidance, and lacks pagination support for list operations. No error classification, retry guidance, or recovery patterns are evident. The descriptions are adequate (50-150 chars) but lack depth around when to use each tool vs. similar ones, and no dependency hints are provided. The output schema for tools is mentioned in docstrings but not formally defined in a way that's parseable by clients.
Analyze content using AI models
Chat with AI models from various providers
Compare responses from multiple AI models
Generate content using AI models
List all available AI models from all configured providers
Input schemas incomplete or not visible for 4 of 5 tools. Only `chat` tool shows full parameter schema with types and descriptions in code. Tools `list_models`, `compare`, `analyze`, and `generate` have parameter definitions inferred from docstrings only, lacking formal JSON Schema with type constraints.
Output schemas not formally documented. Tools return Dict[str, Any] with no documented field structure. Clients cannot determine what fields to expect, forcing them to infer from example calls. This violates the pattern:tool requirement for documented return types.
Error handling lacks recovery guidance. All tools return {'error': str(e)} with no actionable next steps, no error classification (retryable vs user-fixable vs fatal), and no suggestions for self-correction. Pattern review:recovery-guide requires error responses to tell the LLM what to do next.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 44 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 35 | - | v1 |
No pagination support for list_models. The tool can return an unbounded list of models. For servers with many providers and models, this will blow context windows. Should support limit and offset/cursor parameters per pattern:paginated-result.
analyze tool hard-codes analysis prompts without allowing user customization. The prompts dict is fixed for code/text/security/performance/general, with no way for the LLM to provide a custom analysis directive. This limits tool composability and forces tight coupling to predefined analysis types.
chat tool accepts 'stream' parameter but return type differs (async generator vs dict). When stream=True, the implementation collects all chunks into a single string and returns {'content': ..., 'model': ..., 'provider': ...}. When stream=False, it returns response.model_dump(). Output schema is inconsistent and undocumented.
compare tool lacks per-model success/failure tracking. If one model fails, the response includes {'model': model, 'error': str(e)}, but there's no way for the LLM to distinguish which models succeeded and which failed without inspecting the response structure. Should return structured per-item status.