MCP server that gives Claude Code a ChatGPT second opinion. Cross-model review without copy-pasting.
Two tools with clear action verbs, good descriptions (158 - 195 chars, within 10 - 1024 baseline), and complete JSON schemas with typed parameters. Both schemas are present and valid. However, critical gaps exist: (1) ask_chatgpt's 'system' parameter lacks a description entirely; (2) compare_approaches's 'model' parameter has null description; (3) no output schemas documented for either tool, LLMs cannot predict return structure; (4) no error handling guidance ('try X if this fails'), errors are returned but non-actionable; (5) 'temperature' parameter description is minimal ('Creativity level 0-2') and lacks format constraints; (6) no mention of costs, rate limits, or prerequisites (OPENAI_API_KEY must be set). The tools are well-named (ask_chatgpt, compare_approaches) and use enums correctly for 'model' and 'temperature' bounds, but lack the rigor expected for production: no pagination (not applicable here), no batch operations, minimal error recovery guidance.
Ask ChatGPT for a second opinion, alternative approach, or different perspective. Use this to compare thinking, get a baseline to improve on, or learn from how a different model approaches the same problem.
Send the same prompt to ChatGPT and get back its response along with a structured comparison framework. Good for benchmarking your output against another model's output.
Missing parameter descriptions: 'system' param in ask_chatgpt has no description; 'model' param in compare_approaches has null description. LLMs cannot infer when or how to use these parameters.
No output schemas documented. Tool returns are inferred from code (text content with token counts for ask_chatgpt; comparison structure for compare_approaches). LLMs lack clarity on what fields to expect, preventing reliable downstream planning and extraction.
Error handling is non-actionable. Both tools catch errors and return isError: true with a generic 'OpenAI API error: <message>' response. No recovery guidance (e.g., 'Check OPENAI_API_KEY is set' or 'Rate limited, try again in 60s'). LLM cannot self-correct or plan next steps.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 66 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 40 | - | v1 |
No tool-level prerequisites or constraints documented. Missing: requirement for OPENAI_API_KEY environment variable, API rate limits, cost implications, model availability, or when to use each model variant.
'temperature' parameter description is vague ('Creativity level 0-2. Lower = more focused.'). Does not explain valid range enforcement, semantic meaning of 0 vs 1 vs 2, or how LLM should choose it in different contexts.