The server defines 8 tools with mostly complete schemas and descriptions. Naming follows verb_noun patterns (tts_*, stt_*, voice_*) which is clear and action-oriented. However, descriptions are inconsistent in depth, some are quite brief (e.g., tts_stop: 22 chars), and several parameter descriptions lack specificity around constraints, formats, and error recovery. Parameter schemas are present and typed, but missing validation details like string format/length constraints. Output schemas are not documented in the provided code. Error handling is present (evidenced by test cases checking for validation_error, tts_error, file_not_found) but recovery guidance is minimal. The test suite is comprehensive, which suggests solid implementation, but the tool interface itself could be LLM-optimized further.
Tools (8)
stt_list_modelsread onlysource verified77/100
List available STT models, optionally filtered by provider
stt_listenread onlyauthsource verified70/100
Listen via microphone and transcribe speech to text
Document output schemas for all tools. Example: tts_speak should return {success: boolean, text: string, status?: string, error?: string}. tts_list_voices should return {voices: [{name: string, language?: string, platform?: string}]}. This enables LLMs to chain tools correctly.
Add explicit parameter constraints. For tts_speak: speed description should be 'Speed in words per minute (50-400, default from config)'. For stt_listen: duration should be 'Duration in seconds to listen (1-60, default 5)'. For voice_config volume: 'Volume from 0.0 (silent) to 1.0 (max)'.
Expand brief descriptions to 50-150 chars. tts_list_voices: 'List available text-to-speech voices on the current platform. Call this first to see voice options before using tts_speak.' tts_stop: 'Stop currently playing TTS audio. Safe to call even if no audio is playing.'
Add error recovery guidance to descriptions. tts_speak_provider: 'Requires provider API key (OPENAI_API_KEY, ELEVENLABS_API_KEY, etc). If credentials missing, error will indicate which key to set.' stt_transcribe_file: 'Supported formats: WAV, MP3, M4A, FLAC. If file format unsupported, error will specify supported formats.'
Split voice_config into two tools: get_voice_config (returns current settings) and set_voice_config (updates one or more settings). This follows single-responsibility and makes LLM intent clearer.
Add parameter format hints. For tts_speak 'voice' parameter: 'Voice name (e.g. Samantha on macOS, Alex on Linux). Call tts_list_voices first to see available options.' For stt_transcribe_file 'language': 'BCP-47 language code (e.g. en, fr, es, de). If omitted, auto-detects.'
Score history
Overall score trend
↑ 1 points across a rubric change (v1 → v2)
57/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
D
57
2026-07-28+
v2
2026-03-09
D
56
-
v1
voice_config
writesource verified69/100
Get or set voice configuration options
Error recovery guidance missing, error messages in tests (validation_error, tts_error, file_not_found) have no documented recovery steps for LLMs
voice_config has weak naming (generic verb) and unclear semantics around 'action' parameter, should be split into get_voice_config and set_voice_config
voice_config
Document idempotency and side effects. tts_stop: 'Idempotent. Safe to call repeatedly.' tts_speak: 'Has side effect: plays audio immediately or enqueues for playback. Not idempotent if called repeatedly with same text (may play multiple times).'
Add pagination/limit guidance if results can grow large. stt_list_models currently has no limit hint. Consider: 'Returns up to 50 available models. For filtered results, use provider parameter.'
Return contextual error messages. Instead of generic 'validation_error', return 'Invalid provider: got "amazon", must be one of: elevenlabs, openai, google, system'. This lets LLMs self-correct in one retry.
Document platform dependencies. tts_speak works on macOS (say), Linux (espeak), Windows (PowerShell). Document fallback behavior and error cases in description.