The server defines 4 tools with reasonably clear naming and documented input schemas. However, there are significant gaps: parameter descriptions lack depth and guidance, output schemas are undocumented (no structured response format specified), error handling is absent, and composition patterns are weak. Tool names follow verb_noun convention (make_call, quick_call, list_recordings, get_transcript), which is a strength. However, the descriptions, while present, are brief and lack the actionable guidance required for optimal LLM planning. The make_call and quick_call tools exhibit partial redundancy (both make calls, differ only in configuration granularity), suggesting a composition issue. Most critically, there is no documented output schema for any tool, LLMs cannot plan downstream operations when they don't know what fields to expect.
Get the full transcript of a specific call by ID. Use list_recordings first to get call IDs.
List recent call recordings made through this MCP server session. Returns call ID, timestamp, phone number, duration, and transcript.
Make a SIP phone call with TTS prompt and transcription. Returns call duration, transcript, and recording info. Use this for full-featured calls with custom configuration.
Make a quick SIP call with minimal configuration. Simpler than make_call, uses defaults for everything. Perfect for simple test calls or quick messages.
No output schemas documented for any tool. LLMs cannot infer return types, field names, or structure. This prevents downstream tool composition and forces agents to guess at response shape.
Parameter descriptions are minimal and lack actionable guidance. E.g., 'timeout' has no guidance on what happens if a call exceeds it (does it hang, fail gracefully, or retry?). 'phone' should specify accepted formats and error handling for invalid E.164.
No error handling guidance. Tools have no description of failure modes, recovery paths, or error response structure. E.g., what happens if the phone number is invalid, the SIP gateway is down, or TTS fails? Agents have no recovery strategy.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 61 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
make_call and quick_call are redundant. Both make SIP calls; quick_call is just make_call with hardcoded defaults. This violates the single-responsibility principle and forces LLMs to reason about when to use each. Recommend merging into one tool with optional parameters, or clearly document the distinction (e.g., quick_call for <30-char prompts, make_call for advanced use cases).
make_call 'voice' parameter lacks enum constraint. Description lists options (alloy, nova, shimmer, echo, fable, onyx) as free text. Should be declared as an enum in the schema to prevent LLM hallucination of invalid voice names.
list_recordings 'limit' parameter has no guidance on reasonable bounds. Can an agent request limit=999999? Should it? No min/max specified. Recommend adding explicit bounds (e.g., 1 - 100) and documenting the default (10).
get_transcript 'call_id' parameter type is listed as 'number' but should clarify if it's an integer. More importantly, no guidance on what happens if the call_id doesn't exist. Does it return null, an error, an empty string? This ambiguity forces the LLM to guess.
No documented relationship between list_recordings output and get_transcript input. Presumably get_transcript accepts call_id values from list_recordings, but this is not explicitly stated in tool descriptions. Tool chaining documentation missing.