Local MCP voice coach for Codex, Claude, and other MCP clients with English pronunciation, grammar, fluency, phoneme-level feedback, and practice drills
The server demonstrates solid foundational quality with clear verb-based naming, consistent descriptions, and well-defined schemas across 8 tools. Tool descriptions are descriptive (100-300 chars) and answer WHAT/WHEN/WHY questions. All parameters have type constraints and descriptions. However, several patterns from the 54 Agentic Tool Patterns are missing: no explicit error guidance (recovery paths), no pagination for list results, limited composition guidance between tools, and output schemas are not fully documented in the definitions. The use of `tool_annotations` is a positive signal. The server is production-capable for pronunciation training but could be stronger in error handling and result structuring.
Formal pronunciation assessment. Records the user reading a reference sentence, performs word-level alignment, phoneme-diff analysis, forced alignment (if available), and returns a comprehensive assessment with clarity score and detailed coaching feedback.
Free-form voice conversation with pronunciation feedback. Listens for user speech via a voice panel or recorded file, transcribes with Whisper, analyzes prosody, and returns conversational coaching.
Fetch a random practice sentence, optionally filtered by pronunciation focus (th, f_v, r_l, vowels, general) and difficulty (beginner, intermediate, advanced). Returns structured sentence with IPA transcription.
Poll the status of a background voice capture session. Returns the current recording state, transcript (when ready), and assessment results once the recording finishes.
Guided practice drill. Records the user reading a reference sentence aloud, compares against the reference, and provides detailed phoneme-level feedback with word-stress and fluency coaching.
Start a background voice recording session. Returns immediately with a session ID. The recording continues in the background until the specified duration or silence is detected. Use get_voice_status to poll the result.
Output schemas not documented in tool definitions. While input schemas are explicit and typed, return types for all 8 tools are missing from the visible definitions. LLMs cannot infer what fields to expect from converse, practice, assess, etc., forcing them to reason about downstream data structures without guidance.
No error handling or recovery guidance. Tool descriptions do not explain what errors can occur or how the LLM should recover (e.g., if recording fails due to microphone unavailability, what should the agent do?). This violates the recovery-guide pattern.
Tool composition unclear. The relationship between `record` + `get_voice_status` and the direct recording tools (converse, practice, assess) is not well documented. When should an agent use `record` vs calling `converse` directly? Requires implicit multi-step inference.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | B | 70 | <=2025-11-25 | v2 |
Compare a new attempt against a previous recording at the same reference sentence. Tracks improvement in clarity score and issue resolution.
Upload a base64-encoded WAV file captured by a browser voice panel for assessment. Stores it as a temporary WAV and returns the local file path for later reference by assess/practice/converse tools.
Partial parameter validation hints. `duration_seconds` accepts 1 - 120 but descriptions do not explicitly state the range in the parameter description text itself, only in prose. JSON Schema likely enforces this, but LLMs cannot read Schema bounds reliably; they depend on description text.
Missing pagination guidance. `get_sentence` returns a single sentence, but if future versions return multiple results or a sentence bank, there is no pagination or limit specification framework documented.