Multi-model AI gateway for MCP clients — routes prompts to Gemini, OpenAI, Anthropic, xAI, DeepSeek, Moonshot, OpenRouter, and local models with conversation memory
Static source inference · medium confidence · detected: Logging
Deprecated protocol patterns detected
Summary
The Vox MCP server has significant definition quality issues. While it contains three tools with basic descriptions, critical infrastructure for schema validation and parameter documentation is either absent or not visible in the provided source. The naming convention is reasonable ('chat', 'listmodels', 'dump_threads', mostly action-based), but schemas are not explicitly visible in the code sample, parameter descriptions are missing, and tool definitions appear to be inferred rather than explicitly registered with full schema objects. The base_tool.py excerpt shows abstract method signatures for get_input_schema() but no concrete implementations are visible. This is a STDIO-only server with limited introspection capabilities, which compounds schema visibility issues.
No input schemas visible in provided source code. Schema definitions referenced in abstract method get_input_schema() but no concrete implementations shown for any tool.
Tool descriptions are generic and lack specificity. 'Multi-model AI gateway with conversation memory' does not answer WHEN to use chat vs other tools, WHAT parameters it accepts, or WHAT it returns. Baseline is 194 chars; these are 40-50 chars.
Parameter descriptions are missing from visible code. The abstract BaseTool class shows the structure but no concrete parameter documentation is visible for any tool. LLMs cannot infer meaning from param names alone (e.g., 'model', 'provider', 'thread_id').
chatlistmodelsdump_threads
HIGH
Recommendations
Add concrete input schema implementations in each tool class. Return a dict with 'type': 'object', 'properties': {<param>: {'type': <type>, 'description': '<clear description>'}, 'required': [<required_params>]}. Example for chat: {'type': 'object', 'properties': {'model': {'type': 'string', 'description': 'AI model to invoke (e.g., gemini-2.0-flash). Call listmodels() first to see available options.'}, 'prompt': {'type': 'string', 'description': 'User message or prompt (required). Max 128k tokens.'}, 'temperature': {'type': 'number', 'minimum': 0, 'maximum': 2, 'description': 'Sampling temperature. 0=deterministic, 2=max randomness. Default: 0.7.'}}, 'required': ['model', 'prompt']}
Expand tool descriptions to 100-200 characters following pattern:tool-description. Example: 'Send a prompt to any supported AI model (Gemini, OpenAI, Anthropic, DeepSeek, etc.). Maintains conversation memory across turns. Returns model response and usage stats. Call listmodels() first to see available models and pricing.'
For 'chat' tool, rename to 'send_message', 'query_model', or 'invoke_llm' to clarify action verb. If maintaining 'chat', add disambiguator in description: 'Chat with an AI model, not a person. To chat with users, use send_dm.'
Document output schema for each tool. Example for chat: 'Returns {response: string, model: string, tokens_used: {input: int, output: int}, thread_id: string, timestamp: ISO8601}.' Example for listmodels: 'Returns {models: [{name: string, provider: string, pricing: {input_usd_per_mtok: float, output_usd_per_mtok: float}, max_tokens: int}], total_count: int}.'
Spec posture evidence
Inferred effective spec: 2026-07-28+.
Relies on Logging (deprecated) - log to stderr or use OpenTelemetry
'chat' tool name lacks action verb prefix. 'send_message', 'invoke_model', or 'query_model' would be clearer. 'chat' is ambiguous, does it retrieve chat history, send a message, or configure a conversation?
Output schemas are not documented in visible source. LLMs need to know what fields to expect from chat(), listmodels(), and dump_threads() to plan downstream operations and extract data correctly.
No error handling guidance visible. If chat() fails due to invalid model, malformed prompt, or API timeout, LLMs receive no actionable recovery path. Error responses should state: 'Model X not found. Available models: [list]' or 'Prompt exceeds token limit (max: Y). Reduce content by Z%.''
Tool definitions inferred from class structure; explicit registration with full schema and description objects not visible. Concrete tool registration (e.g., call to register_tool() with schema dict) not shown.
chatlistmodelsdump_threads
Add parameter constraints to schemas. For chat/model: use enum of available models or accept string + validate server-side. For temperature: constrain to 0 - 2. For prompt: set minLength=1, maxLength=128000. For thread_id: describe as optional UUID or alphanumeric ID; state default creates new thread.
Implement error handling with recovery guidance. Example: if model not found, return '{error: "Model not found", suggested_models: [list of closest matches], message: "Try one of: " + list}.' If prompt too long, return '{error: "Prompt exceeds token limit", max_tokens: 128000, your_tokens: X, message: "Reduce by at least Y tokens."}'
Add tool annotations (from features.toolAnnotations=true). Specify readOnlyHint=true for listmodels and dump_threads (safe to retry). For chat, specify destructiveHint=false, idempotentHint=false if conversation history is mutable (retrying same prompt may get different response if temperature > 0).
For dump_threads, clarify whether this is a debug/export tool or a production-facing feature. If debug, add warning in description: 'Exports active conversation state as JSON. Use for debugging/migration only; not for production sync.' If production-facing, document structure of returned JSON and any rate limits.
Accept human-friendly identifiers for models. Instead of requiring 'gemini-2-0-flash', accept 'Gemini 2.0' and resolve internally. Document accepted formats in parameter description.
Add pagination to listmodels if result set is large. Include limit (default 20, max 100), offset/page_number, and total_count in response. Example: 'Returns {models: [...], total_count: 47, limit: 20, offset: 0, has_more: true}.'
Implement conversation memory as transparent parameter. Add optional 'thread_id' param to chat; state: 'Unique conversation identifier. Omit to start new thread. Include to continue prior conversation. Returns thread_id for chaining.' Return thread_id in response so agent can reuse it.
Document the MCP_PROMPT_SIZE_LIMIT from config in tool schema. Add to chat description: 'Prompt limited to {limit} tokens per MCP spec. Very large prompts will be truncated.' Make limit visible as a constraint, not a surprise failure.