MCP server that provides consulting agents (Darren, Sonny, Sergey, Gemma) using various LLM models with specialized capabilities including reasoning, extended thinking, web search, and large-context analysis
This server defines 3 tools (consult_darren, consult_sonny, consult_sergey) with minimal but present parameter schemas. All tools have descriptions and documented input parameters, but descriptions are sparse (60-90 chars, below the 194-char baseline for A+ tools), parameter descriptions lack depth regarding constraints and error cases, and no output schemas are documented. The code shows timeout handling and error logging, but error messages are not actionable for LLMs. The tools themselves represent a reasonable single responsibility (consult with a specific expert model), but descriptions do not explain when to choose one expert over another or what each is optimized for. Parameter validation is present (API key verification) but not comprehensive. Output parsing is complex and fragile, suggesting the actual response schema is not well understood or documented. Overall, this reads like a functional prototype rather than a production-grade tool suite.
Consult with Darren using OpenAI's responses API with o3-mini and high reasoning.
Consult with Sergey using OpenAI's Responses API with GPT-4o and web search.
Consult with Sonny using the Anthropic API with extended thinking enabled.
Tool descriptions are too terse (60 - 90 characters). Baseline for A+ tools is 194 chars. Missing context on WHEN to use each consultant (e.g., 'Darren for coding/debugging with deep reasoning', 'Sonny for extended thinking', 'Sergey for web-search-backed answers'). LLMs cannot disambiguate between the three without more actionable differentiation.
Output schema is NOT documented. Code shows complex response parsing logic (checking for 'output' arrays, fallbacks to 'output_text', handling 'content' blocks in Anthropic responses), but the tool definitions do not specify what fields the LLM can expect. LLMs cannot plan downstream use of the output.
Parameter descriptions lack constraint details. The 'prompt' parameter in all three tools has only a brief description ('The prompt/question to send to...') but no guidance on length limits, required structure, or format. The 'search_query' parameter in consult_sergey notes it is optional but does not explain the fallback behavior if omitted.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | <=2025-11-25 | v2 |
Error messages are not LLM-actionable. When API requests fail (e.g., 'OpenAI API request failed: <str(e)>'), the LLM receives a generic error with no recovery guidance. Missing: (1) classification (retryable vs fatal), (2) suggested next steps (e.g., 'Check API key, retry', 'Contact support'), (3) specific constraint violations.
Tool names use 'consult_<name>' which is clear but does not follow the <verb>_<noun> pattern dominant in production tools (get, create, search, send). 'consult_' is domain-specific jargon. More neutral naming would be 'ask_darren', 'ask_sonny', 'ask_sergey' or similar. However, 'consult_' is acceptable if the server is meant as a specialized domain tool.
No distinction between the three tools in descriptions regarding model capabilities. The descriptions state WHAT API is used (OpenAI Responses, Anthropic extended thinking, OpenAI with web search) but not WHY an LLM should pick one consultant over another. Missing: guidance like 'Use Darren for deep reasoning on code problems', 'Use Sonny for extended internal reasoning', 'Use Sergey for answers requiring current web information'.