TTS MCP server and CLI for language learning (ElevenLabs, AWS Polly, OpenAI)
The server defines 2 tools with explicit schemas and descriptions visible in the source code. Both tools follow a clear verb_noun pattern (synthesize, synthesize_batch) appropriate for their domain. Descriptions are comprehensive (250-350 chars range) and explain WHAT the tools do, WHEN to use them, and what they return. All input parameters have types and descriptions. However, there are significant gaps: no output schema documentation for return types, no error handling guidance, no parameter constraints (enums for language codes), and parameters accept arbitrary string values where validation is needed. The tools are well-designed for composition (one for single text, one for batch), but lack defensive measures like range validation, dry-run modes for file operations, and clear error recovery paths.
Synthesize text to an MP3 audio file. Args: text: The text to convert to speech. With ElevenLabs eleven_v3, you can embed audio tags in square brackets anywhere in the text to control delivery — e.g. [tired], [excited], [whisper], [sad], [sigh], [laughs], [dramatic tone]. Tags are free-form; the model interprets them as performance cues. Combine with punctuation (ellipsis for pauses, ! for emphasis) for best results. Tags only work with ElevenLabs eleven_v3 model. voice: Voice name. Default: provider's default voice (currently matilda for ElevenLabs, joanna for Polly, nova for OpenAI). If language is provided without voice, a suitable default voice for that language is selected automatically. language: ISO 639-1 language code (e.g. 'de', 'ko', 'fr'). Enables language-aware voice selection and validation. With Polly, validates voice-language compatibility. With ElevenLabs/OpenAI, passed through (voices are multilingual). rate: Speech rate as percentage (90 = 90% speed, good for language learners). Defaults to 90. ElevenLabs ignores rate; use audio tags like [rushed] or [drawn out] instead. auto_play: Open the file in the default audio player after synthesis. Defaults to true. output_path: Full path for the output file. If not provided, a file is auto-generated in output_dir. output_dir: Directory for output. Defaults to TTS_OUTPUT_DIR env var or ~/langlearn-audio/. stability: ElevenLabs voice stability (0.0-1.0). Ignored by other providers. Defaults to provider default. similarity: ElevenLabs voice similarity boost (0.0-1.0). Ignored by other providers. Defaults to provider default. style: ElevenLabs voice style/expressiveness (0.0-1.0). Ignored by other providers. Defaults to provider default. speaker_boost: ElevenLabs speaker boost toggle. Ignored by other providers. Defaults to provider default. Returns: JSON string with path, text, voice, and language fields.
Synthesize multiple texts to MP3 files. Args: texts: List of texts to synthesize. With ElevenLabs eleven_v3, embed audio tags like [tired], [excited], [whisper] in text. voice: Voice name for all texts. Default: provider's default voice (currently matilda for ElevenLabs, joanna for Polly, nova for OpenAI). If language is provided without voice, auto-selects. language: ISO 639-1 language code (e.g. 'de', 'ko'). rate: Speech rate as percentage. Defaults to 90. merge: If true, produce one merged file instead of separate files per text. Defaults to false. pause_ms: Pause between segments in milliseconds when merging. Defaults to 500. auto_play: Open the file(s) in the default audio player after synthesis. Defaults to true. output_dir: Directory for output files. Defaults to TTS_OUTPUT_DIR env var or ~/langlearn-audio/. stability: ElevenLabs voice stability (0.0-1.0). similarity: ElevenLabs voice similarity boost (0.0-1.0). style: ElevenLabs voice style/expressiveness (0.0-1.0). speaker_boost: ElevenLabs speaker boost toggle. Returns: JSON string with list of results, each containing path, text, voice, and language fields.
No output schema documentation. Tool descriptions state return format ('JSON string with path, text, voice, and language fields') but no structured schema is provided for downstream tool composition.
Parameters accept arbitrary string values where validation is needed. Examples: 'voice' has no enum of valid voice names; 'language' accepts any string instead of ISO 639-1 enum; 'rate' is unbounded integer (docs say percentage but code allows any value).
No error handling guidance. Descriptions do not explain what happens on failure (invalid voice, missing API credentials, file write failure, network timeout). LLM receives no recovery hints.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 63 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Floating-point parameters (stability, similarity, style) lack explicit range constraints in JSON Schema (minValue/maxValue). While descriptions mention 0.0-1.0, the schema does not enforce it, allowing LLMs to pass invalid values.
No confirmation/dry-run pattern for destructive operations. The 'auto_play' parameter auto-executes file operations (creates audio files, opens players) without user confirmation. High-risk for accidental side effects.
Parameter relationships undocumented. 'stability', 'similarity', 'style', 'speaker_boost' are ElevenLabs-specific and ignored by other providers, but this dependency is buried in the description. LLM will not understand why some params are silently ignored.