USB-only PC gateway and local MCP server for cellular calling. No SIP/RTP/Asterisk/STUN/LAN/Wi-Fi/WSS.
AgentCall defines 11 tools with complete JSON schemas and descriptions. Naming follows verb_noun convention (dial, answer, reject, hangup, speak, send_dtmf). Descriptions are present (avg ~100 chars) but lack depth on WHEN to use each tool and dependencies between them. All parameters have type definitions and descriptions with format constraints (e.g., E.164 for phone, 1-128 chars for IDs). However, output schemas are not documented, callers cannot see what dial() or speak() return. Error handling is absent: no guidance on retryability, recovery steps, or what happens on network failure. Idempotency keys are present on write operations (good pattern), but no explicit documentation of idempotent semantics. Tool composition is sound (each does one thing), but descriptions don't explain call sequencing (e.g., must wait_for_incoming_call before answer).
Answer an incoming call
Query MCP server capabilities including supported tools, transport, protocol version, framing, and policy
Initiate an outgoing call to a destination
Terminate an active call
Prepare a speech response for the current call
Reject an incoming call
Send DTMF tones during an active call
Output schemas not documented. Callers cannot see what dial(), speak(), or wait_for_turn() return. This forces LLMs to guess at response structure and breaks tool chaining.
Descriptions lack WHEN-to-use guidance and call sequencing. E.g., wait_for_incoming_call description doesn't explain it must be called before answer(). Agents cannot infer prerequisites.
No error handling guidance. Descriptions don't explain what happens on timeout, network failure, or invalid call state. LLMs cannot determine if errors are retryable.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | D | 59 | 2026-07-28+ | v2 |
Speak text during an active call
Query gateway status including device connection, authentication, recording state, realtime provider status, and current call information
Wait for an incoming call with optional timeout
Wait for a turn in an active call with optional timeout
prepare_speech vs speak naming ambiguity. Both send audio to a call. Descriptions don't clarify when to use prepare_speech (pre-recording?) vs speak (immediate). LLMs may conflate them.
Idempotency keys present but semantics undocumented. Descriptions don't explain what happens if the same idempotencyKey is reused, or how long deduplication lasts.