Text-to-speech MCP server supporting macOS Say, Google Gemini TTS, and OpenAI TTS with audio playback and file saving capabilities
The server implements a single TTS tool 'speak' with a structured JSON Schema that includes input types and descriptions. However, the tool definition reveals critical gaps: the description is adequate (103 chars, within 10-1024 range), but several parameters lack clear guidance on constraints, expected formats, and interdependencies. The schema is visible and typed, but parameter descriptions are terse and do not explain when to use which provider or voice, nor do they document mutual exclusivity (e.g., provider-specific parameters like 'stability' and 'similarity_boost' that only apply to ElevenLabs). Error handling documentation is absent from the source provided, no recovery guidance for invalid provider, missing voice IDs, or API failures. The tool performs a write operation (speaks/saves audio) but lacks documentation about idempotency, confirmation patterns, or file-save side effects.
Convert text to speech using the specified TTS provider and play the audio or save it to a file
Parameter descriptions lack format and constraint guidance. 'voice' parameter is described as 'Voice identifier for the selected provider' but does not specify valid voice names per provider, expected format, or how to discover available voices.
Undocumented parameter dependencies. Parameters 'stability' and 'similarity_boost' are marked 'ElevenLabs compatible' but tool description does not mention ElevenLabs as a provider option, nor does it document which parameters apply to which provider (say vs google vs openai vs elevenLabs).
Missing output schema documentation. The tool description does not specify what the tool returns on success (e.g., file path if saved, confirmation message, error details if playback fails). LLMs cannot plan follow-up actions without knowing the response structure.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-21 | F | 49 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 38 | - | v1 |
No error recovery guidance. The tool description does not explain what to do if the provider is unreachable, voice ID is invalid, or file save fails. Error messages should tell the LLM what to do next, not just report failure.
Parameter 'model' is described as 'Model identifier for cloud providers (google, openai)' but the enum constraint lists 'say, google, openai', 'say' is a macOS system tool, not a cloud provider with a model field. This inconsistency can confuse LLMs about when 'model' is required.
No documentation of write operation semantics. The tool description does not clarify that 'speak' is a write operation (plays audio or saves to disk), making side effects ambiguous. Agents need to know if the call is idempotent and whether it can be safely retried.
Missing provider discovery tool. Users and LLMs do not have a way to discover available voices, models, or providers without trial and error. A 'list_voices' or 'describe_provider' discovery tool would prevent invalid calls.