MCP server for making AI-powered phone calls with automatic instruction generation using OpenAI's o3 model. Provides two tools for simple and advanced call control with SIP-based VoIP integration and G.722 wideband audio support.
The server defines 2 tools with comprehensive schemas and detailed descriptions. Both tools have complete JSON Schema definitions with proper typing, enums, and constraints. Descriptions are lengthy and contextual, explaining the o3 model integration. However, there are significant gaps in error handling guidance, no output schema documentation, and the tools combine multiple responsibilities (instruction generation + call execution) that could be split. The naming follows verb_noun convention ('simple_call', 'advanced_call') but lacks clarity on what distinguishes them until the descriptions are read. Parameter descriptions are generally good but some lack actionable guidance for LLM error recovery.
Make an AI-powered phone call with granular control over all call brief components. Specify each field individually instead of using a natural language brief. The system will use o3 to generate optimized instructions from your structured data. This tool gives you fine-grained control over every aspect of the call brief - target name, goal, constraints, fallback options, formality level, industry context, etc. The o3 model will still generate sophisticated instructions, but based on your structured input rather than a free-form brief. Use this when you need precise control over specific parameters or have complex requirements that benefit from structured specification.
Make an AI-powered phone call with automatic instruction generation. Requires a brief description, your name, and phone number. The system will use OpenAI's o3 model to generate detailed instructions. Why this works better than manual instructions: OpenAI's real-time voice models are optimized for speed, not sophistication. They struggle with complex, goal-oriented tasks without very specific instructions. The o3 model automatically transforms your simple brief (like 'Call the restaurant and book a table for 2 at 7pm') into sophisticated, detailed instructions with conversation states, fallback strategies, and appropriate tone - saving you from writing lengthy manual instructions while achieving much better call results. IMPORTANT: for the brief, use the language the call will be held in - the language of the call will be inferred from the brief!
No output schema documented. The tools return CallResult but the structure, fields, and types are not defined in the tool registration. LLMs cannot plan downstream actions or extract call outcomes without knowing the response shape.
No error handling guidance. Tool descriptions do not explain what happens on failure (invalid phone number, network error, OpenAI API failure, SIP connection loss) or what the LLM should do next (retry, ask user, fallback). Agents will be stranded on error.
Tool names are ambiguous. 'simple_call' vs 'advanced_call' does not convey actionable intent, the distinction is only clear after reading long descriptions. Better names: 'call_with_auto_instructions' (uses o3 generation) vs 'call_with_custom_brief' (structured params). Current names force LLM to read full docstrings to choose between them.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 56 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Example values embedded in descriptions. The 'brief' parameter includes example text ('Call Bocca di Bacco restaurant...') which LLMs frequently reuse literally rather than adapting to context. Replace with format/length constraints in the schema.
Destructive operation not flagged. Both tools initiate phone calls (side effect: outbound call, potential charges, state modification) but lack a destructiveHint annotation. LLMs need explicit marking to treat these as irreversible and require confirmation.
Parameter interdependencies underdocumented. For 'advanced_call', fields like 'date', 'time', 'location', 'constraints' are semantically linked but described independently. No guidance on which combinations are valid or how the o3 model weighs conflicting inputs.
No pagination or output limits documented. CallResult structure is unknown, but if it returns full call transcripts or event logs, there is no mention of limits, truncation, or streaming. Large calls could exceed context windows.
'config_path' parameter in 'simple_call' defaults to 'config.json' but no guidance on where this file is searched (cwd, home dir, server root). LLMs cannot reason about file paths reliably.