MCP server for VOICEVOX text-to-speech integration
Single tool with complete JSON Schema and basic descriptions. Tool name 'speak' is verb-led and clear. Schema includes proper constraints (speedScale: 0.5-2.0, volumeScale: 0-2.0, required params). However, descriptions are in Japanese and relatively brief (under 100 chars for most params). No output schema documented. Error handling exists but returns plain text instead of structured guidance. Tool does one thing well but lacks production-grade documentation polish and LLM-optimization for English-language agents.
VOICEVOXを使用してテキストを読み上げます
Descriptions are in Japanese; non-English agents cannot reliably interpret parameters. Tool descriptions and param docs should be in English or provide bilingual coverage.
No output schema documented. Callers do not know what 'speak' returns beyond a text string. What fields are in the response? Is 'おしゃべり完了' (conversation complete) always the success message? Output structure must be formalized.
Error responses return unstructured text (e.g. 'エラー: <message>'). LLM has no guidance on recovery. Should classify errors as retryable (network timeout) vs. user-fixable (invalid speaker ID) and provide actionable next steps.
Parameter 'speaker' lacks validation context. Description says 'VOICEVOXの話者番号' (speaker number) but does not specify valid range, how to discover valid IDs, or what happens if an invalid speaker is passed. LLM cannot self-correct.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 58 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 43 | - | v1 |
Tool modifies state (plays audio, has side effects) but descriptions do not explicitly state this is a command/action tool with irreversible consequences. Agents need to know they cannot safely retry this call.