This OpenAI MCP server has moderate definition quality but suffers from several critical gaps. Tool naming follows verb_noun conventions (get_*, query_*) which is good, but descriptions are minimal and lack context. Input schemas are present but incomplete, parameters lack descriptions in critical places, and parameter validation is weak. Output schemas are entirely undocumented. Error handling returns text responses rather than structured guidance. The server exposes a non-standard model default ('gpt-5.6-sol') which appears to be placeholder/fictional. Overall, definitions are functional but fall short of production-grade quality expected for agent tool composition.
Get information about an OpenAI model
Get the name and version of this MCP server
Query OpenAI API with a prompt and get a response
Tool descriptions are minimal (10-30 chars) and lack actionable context. 'Query OpenAI API with a prompt and get a response' does not explain when to use this vs alternatives, what the response structure is, or prerequisites. LLMs cannot reliably select this tool without richer description.
Output schemas are completely undocumented. Tools return structured text responses but callers do not know: what JSON structure is included, what fields are guaranteed, what happens on error, or how to chain outputs to downstream tools. This violates the requirement that LLMs know what to expect from tool results.
Parameter 'reasoning_effort' has an enum constraint but the enum includes 'none' which is not a valid effort level per OpenAI docs. Enum should be ['low', 'medium', 'high', 'xhigh', 'max']. The 'none' value will cause runtime errors.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 40 | - | v1 |
Default model 'gpt-5.6-sol' appears to be non-existent or placeholder. This will cause every call with default params to fail. Either use a real model (gpt-4o, gpt-4-turbo, etc.) or require the parameter.
Parameter descriptions are missing or vague. 'max_completion_tokens' says it is for 'GPT-5.x models' but does not explain when to use it vs 'max_tokens', or whether both can be specified. 'reasoning_effort' lacks any explanation of what each level does or when to use 'low' vs 'max'. LLMs cannot reason about these dependencies.
Error handling returns plain text error messages without structured guidance. When 'Error querying OpenAI: 401 Unauthorized' is returned, the LLM has no way to know if the error is retryable, user-fixable (e.g. provide API key), or fatal. Recovery paths are not documented.
OpenAI API key is loaded from config.json and passed to the OpenAI client via server-side injection, which is correct. However, the config loading does not validate that the key is present before initializing; it exits with a generic error message. If the key is invalid, the server will start but fail silently on first use.
The 'prompt' parameter in query_openai has no description of expected format, length, or constraints. Should document: 'A string containing the user prompt to send to OpenAI. No length limit enforced client-side; OpenAI API enforces token limits per model.'
Tool 'get_model_info' returns a hardcoded subset of model fields (id, object, created, owned_by) but does not document this. The description should state: 'Returns basic model metadata including ID, object type, creation date, and owner. Does not return pricing, context window, or capability details.'