A simple MCP (Message Control Protocol) server that provides a unified interface to various LLM providers (OpenAI, Anthropic, Google, DeepSeek) using Pydantic AI.
Single tool server with moderate quality. The run_llm tool has a complete input schema with all parameters typed and described, meeting baseline requirements. However, the tool name 'run_llm' is generic and violates verb-noun clarity patterns, 'run' is too vague and doesn't convey the specific action. The tool description is adequate but brief (47 chars), below the 50-200 char LLM-optimized range. Most critically, the output schema is undocumented, the tool returns a JSON string containing LLMResponse fields, but no schema is declared for downstream tool consumers. Error handling is absent; the tool will fail silently on API errors without recovery guidance. The system_prompt parameter has a poor default (empty string) that could cause unintuitive LLM behavior. No rate limiting, secret validation, or error categorization is present.
Run a prompt through an LLM and return the response.
Tool name 'run_llm' is generic and violates verb-noun convention. 'run' and 'do' are discouraged as they lack specificity. Should be 'call_llm', 'query_llm', or 'invoke_llm' to be more explicit about the action.
Output schema is undocumented. Tool returns LLMResponse as JSON string, but the response structure (content, model_name, usage, temperature fields) is not declared. Downstream tools and LLMs cannot predict what fields will be available. This violates the pattern:response-shaper requirement.
No error handling or recovery guidance. If an API call fails (invalid model, rate limit, authentication error), the tool will raise an exception without context for the LLM. Should return structured errors like 'Model not found. Available models: ...' or 'API rate limit exceeded. Retry in 60s.'
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-21 | F | 45 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 54 | - | v1 |
Tool description is only 47 characters, below the 50-200 char LLM-optimized baseline. Missing context on when/why to use this tool vs. calling LLMs directly, what 'unified interface' means, or use cases.
system_prompt parameter defaults to empty string, which provides no system guidance to the LLM. This can produce unintuitive behavior, the LLM operates without behavioral constraints. Should document why empty is the default or suggest a sensible fallback.
No input validation or constraint documentation. temperature parameter should validate 0.0-1.0 range; max_tokens should validate >0 and reasonable bounds (e.g. 1-128000). Currently relies on pydantic to catch, but no guidance in description.
No rate limiting or quota management. An agent in a loop could invoke this tool thousands of times per minute, exhausting API budgets. Should implement rate limiting with backoff guidance.