A Model Context Protocol server for text-to-speech, audio processing, and audio-to-text transcription with support for multiple TTS providers including ElevenLabs and Replicate models
The server provides 10 tools with reasonable naming conventions and documented input schemas. However, there are significant gaps in output schema documentation, error handling guidance, and some descriptions lack LLM-optimization depth. Most tools follow verb-noun naming (text-to-speech, get-audio-metadata, trim-audio) which is good, but descriptions are generic and don't clearly explain WHEN to use each tool versus alternatives. Parameter descriptions are present but often minimal. No output schemas are documented in the source, forcing LLMs to infer return types. Missing error recovery guidance and confirmation mechanisms for destructive operations.
Adjust the volume of an audio file
Transcribe audio files to text with word-level timestamps using WhisperX. Automatically saves transcript with timestamps, sentence-level timestamps, and returns directory reference
Concatenate multiple audio files into one
Convert audio file to a different format
Convert text to speech using ElevenLabs API with character-level timestamps. Automatically saves audio file and timestamp files (character, word, and sentence level) to specified or default audio directory
Extract metadata from audio files including duration, format, bitrate, and other properties
Output schemas are completely undocumented. No tool explicitly documents what fields, types, or structure LLMs should expect in responses. This forces LLMs to guess return types and breaks tool composition, agents cannot reliably chain tools when downstream parameters are unknown.
No error handling guidance in any tool descriptions. Descriptions do not explain what happens on failure, whether errors are retryable, or what the LLM should do next. This violates the recovery-guide pattern.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 48 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 46 | - | v1 |
Get comprehensive information about available TTS models including parameters, capabilities, and characteristics
Split audio file into segments of specified duration
Convert text to speech using Replicate models with support for various TTS models including Chatterbox, Minimax, and others. Automatically saves audio file and creates optional transcript files
Trim audio file to specified start and end times
Descriptions are too generic and do not explain WHEN to use each tool versus similar alternatives. Example: 'Trim audio file to specified start and end times' tells WHAT it does but not WHEN to call it (after editing? after silence detection?). Baseline for LLM-optimized descriptions is 50-200 chars with context; most here are 30-100 chars with minimal guidance.
Two write-heavy TTS tools (text-to-speech, elevenlabs-text-to-speech) lack confirmation or dry-run mechanisms before generating expensive API calls. Agents cannot preview results or confirm direction before incurring costs.
Parameter descriptions for time formats (trim-audio, split-audio) mention 'HH:MM:SS or seconds' but do not validate or guide LLMs on format choice. No enum or pattern constraint. LLMs may pass invalid formats like '1:30' instead of '00:01:30'.
elevenlabs-text-to-speech exposes 'voice' parameter as enum but also accepts 'voice_id' directly. This dual-input pattern is underdocumented, no explanation of when to use each, or whether they're mutually exclusive. Forces LLM to guess or guess wrong.
Numeric parameters (stability, similarity_boost, style, seed) in elevenlabs-text-to-speech lack min/max constraints in descriptions. Parameter descriptions state defaults (0.5, 0.8, 0.0) but no range guidance. LLMs may pass out-of-bounds values like 5.0 for stability.