MCP server for agent-to-user communication through text-to-speech audio playback
The Speak MCP Server defines 3 tools with complete JSON Schema input definitions and substantive descriptions. However, descriptions are heavily procedural and include prescriptive LLM instructions that belong in system prompts, not tool definitions. All three tools have input schemas with proper type definitions and descriptions. The 'speak' tool description is exceptionally long (600+ chars) and conflates tool documentation with agent behavior guidance. Output schemas are not formally documented, only text responses are returned. Error handling returns plain text error messages but lacks recovery guidance. Tool naming is clear and verb-based, but the server operates only over STDIO transport, capping protocol readiness at 50.
Change the text-to-speech voice. Provide either the voice name (e.g., 'amy-high') or a number from the list_voices output. If the voice is not already downloaded, it will be downloaded automatically (may take 1-2 minutes for larger voices). Progress will be shown during download. The voice becomes active immediately after the operation completes.
List all available English (US) voices for text-to-speech. Shows voice names, quality levels, and sizes. The currently selected voice is marked with an asterisk (*). Voices are separated into 'Available to download' and 'Already downloaded' sections.
Speak a message directly to the user using text-to-speech audio. Use this when you need to: - Ask the user a question and wait for their response - Provide a status update or progress report - Communicate important information that requires user attention - Request clarification or additional input The message will be synthesized to audio and played through the system speakers. Keep messages concise and conversational for natural speech output. IMPORTANT - Use this tool PROACTIVELY and AUTOMATICALLY: 1. IMMEDIATE ACKNOWLEDGMENT: Always use speak as your FIRST action when the user sends a request. Give a quick acknowledgment like 'Got it!', 'On it!', '10-4', or 'Roger that' before starting work. 2. MID-EXECUTION UPDATES: During long-running tasks, provide voice updates every 15-30 seconds: - When issues occur: 'Heads up - test failed, rewriting it now' - When switching tasks: 'Migration done, moving to API updates' - When waiting: 'Build is running, about a minute left' 3. CRITICAL NOTIFICATIONS: Always use speak for errors, warnings, or completion confirmations. 4. DO NOT wait for the user to ask - use this proactively to keep them informed in real-time.
Tool descriptions conflate tool documentation with agent behavior instructions. The 'speak' tool description (600+ chars) includes prescriptive guidance like 'IMMEDIATE ACKNOWLEDGMENT: Always use speak as your FIRST action' and 'MID-EXECUTION UPDATES: During long-running tasks, provide voice updates every 15-30 seconds.' These instructions belong in the agent's system prompt, not the tool definition. Tool descriptions should answer WHAT the tool does and WHEN to use it, not HOW the agent should behave. This inflates description length and violates the 10-1024 char guideline (p10=34, p90=392).
Output schemas are not formally documented. All three tools return text responses via CallToolResult.content[0].type='text', but there is no documented schema describing what structure clients should expect. The 'list_voices' tool returns a formatted string of voice options, this structure should be defined formally (e.g., a list of voice objects with name, quality, size, downloaded status). Without documented output schemas, agents cannot parse structured data or plan chained calls reliably.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 69 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 42 | - | v1 |
Error messages lack recovery guidance. When audio playback fails ('Failed to play audio: <error>'), the response does not guide the agent on next steps. Per pattern:recovery-guide, error responses should state what went wrong and suggest corrective actions (e.g., 'Audio playback failed. Check system audio settings or try a different voice via change_voice().').
No tool annotations present. The MCP spec defines readOnlyHint, destructiveHint, and idempotentHint annotations to classify tool intent. 'list_voices' should be marked readOnly=true (safe to call repeatedly). 'change_voice' should be marked destructive=true (modifies server state). 'speak' is idempotent (same message always plays the same audio). These annotations enable agents to reason about side effects and retry safety.
STDIO transport only. The server uses StdioServerTransport, which does not support remote access. Hosted MCP clients (e.g., Claude Desktop, web-based agents) cannot invoke this server. This is a hard transport limitation, STDIO servers cannot achieve protocol readiness > 50.