MCP Server for integrating with Typecast AI text-to-speech API, providing tools for voice generation, voice listing, custom voice cloning, and voice recommendations
The server has 10 tools with mostly complete schemas and descriptions, but exhibits several quality gaps. All tools have descriptions (10-100+ chars) and input schemas with type definitions. Naming is action-oriented (search_, list_, get_, create_, delete_). However, parameter descriptions are sparse or generic, output schemas are not documented, error handling is minimal, and there are no tool annotations (readOnlyHint/destructiveHint). The server uses context variables to inject API keys (good security), but lacks guidance on destructive operations and proper error recovery. Average tool score is 62, placing this in the 'Fair' range with noticeable gaps.
Delete a custom cloned voice
Get details about a specific cloned voice
Get detailed information about a specific voice by voice_id
Create a custom cloned voice from an audio sample
List all custom cloned voices for the authenticated user
List available Typecast voices with optional filtering by model, gender, and age
Play audio file on the local system (only available in local/self-hosted mode)
Output schemas are not documented. The tools return complex objects (voice details, audio data, cloned voice info) but their response structures are not specified. LLMs cannot plan downstream tool calls or extract needed fields without knowing what the response contains.
Missing tool annotations. No tools declare readOnlyHint, destructiveHint, or idempotentHint. This is critical for destructive tools like delete_cloned_voice, the agent cannot infer whether the operation is reversible or requires confirmation.
Parameter descriptions are sparse or missing constraints. Many parameters lack format/range documentation: 'count' in recommend_voices has no min/max bounds; 'audio_pitch' states range but not how the API responds to out-of-range values; 'emotion_intensity' has no guidance on what 0.0 vs 2.0 sounds like.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 66 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 53 | - | v1 |
Get voice recommendations based on a text description of desired voice characteristics
Search the Typecast API documentation
Generate speech audio from text using a specified voice and model
list_cloned_voices has no input parameters but the description does not explain pagination behavior. If results can exceed 50 items, the tool lacks limit/offset/cursor support, forcing the agent to retrieve everything or fail silently on truncated results.
Error handling is implicit. The source shows ToolError exceptions but no guidance on recovery. For example, if text_to_speech fails due to invalid emotion_preset, the LLM receives the error but has no hint to call list_voices or re-try with a different preset.
delete_cloned_voice is destructive but lacks a confirmation/dry-run pattern. The description does not warn the agent that this call is irreversible. No protect_before_delete or soft-delete option.
play_audio accepts a local file path but the description does not explain the security implications or constraints on file access. No mention of sandbox boundaries or whether /etc/passwd can be passed.