A Model Context Protocol server that generates speech audio from text using Microsoft Edge TTS. Supports multi-role conversations, multiple voices, and audio merging.
This server has 2 tools with significant quality gaps. Tool naming is reasonable (verb-noun pattern), but descriptions lack depth and parameter documentation is incomplete. The `text_to_speech` tool has a complex schema with nested arrays, but parameter descriptions are minimal and lack guidance on defaults, constraints, and LLM usage patterns. The `list_voices` tool is minimal but adequate. No output schemas are documented, making it impossible for LLMs to know what structure to expect. Error handling is not evidenced in the visible code. Overall, this is a below-average implementation that would require substantial work for production use.
List all available voices for text-to-speech synthesis.
Generate speech audio from text. Supports multi-role conversations and audio merging.
No output schemas documented for any tool. LLMs cannot predict return structure (file paths? URLs? raw audio?). For text_to_speech with merge_output=true vs false, the response format is completely unclear.
text_to_speech description is generic (56 chars: 'Generate speech audio from text. Supports multi-role conversations and audio merging.'). Does not explain WHEN to use it, what merge_output does, what the return value is, or whether it modifies state. Lacks guidance for LLM selection.
Required parameters in text_to_speech items array (character_name, pause_ms) lack descriptions. character_name says 'Optional character name for dialog context' but is marked required, contradictory. pause_ms has no guidance on valid range or units (clearly milliseconds from name, but not from description).
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 37 | - | v1 |
Speed and pitch parameters use enums (normal/slow/fast, default/low/high) with minimal descriptions. No guidance on what each level means, whether they compound with merge_output, or what defaults apply if omitted.
list_voices description is only 46 chars ('List all available voices for text-to-speech synthesis.'). Does not explain what data structure is returned, whether it includes voice IDs (required by text_to_speech), or sample output. Generic and unhelpful for LLM selection.
No error handling guidance visible in tool definitions. If merge_output fails, if a voice ID is invalid, or if text is too long, what does the tool return? How should an LLM recover? Pattern violation: recovery-guide.