The server implements 6 tools for voice synthesis and cloning via the Waves API. Tool naming follows action-verb conventions (create, get, delete, synthesize), which is good. However, there are significant gaps in schema completeness, parameter descriptions, and output documentation. The `file` parameter in createClone is defined as type 'object' with a generic description rather than proper JSON Schema with nested field definitions. Several tools lack comprehensive output schema documentation. Parameter descriptions exist but are often generic or incomplete, e.g., 'The model to use for synthesis' doesn't specify valid model values or constraints. The server logs errors well internally but does not expose structured error recovery guidance to agents. The definitions are present and functional but fall short of production-grade clarity expected for LLM-driven composition.
Creates a new voice clone on the Waves API using the provided audio file.
Deletes a cloned voice from the Waves API for a specified model.
Fetches the list of cloned voices for a specified model from the Waves API.
Fetches the list of available voices from the Waves API.
Synthesizes speech from text using the Waves API Lightning model.
Synthesizes speech from text using the Waves API Lightning-Large model.
File parameter in createClone lacks proper JSON Schema definition. Defined as generic 'object' type with inline description of nested fields ('content', 'name', 'type') rather than a structured schema with properties, required fields, and type definitions for each nested field.
Model parameter across multiple tools (getClones, deleteClone, synthesizeSpeech) lacks enum constraint. Should declare allowed values as enum ['lightning', 'lightning-large'] rather than free-form string, preventing LLM hallucination of invalid model names.
Output schemas not documented for any tool. No description of what getVoices, getClones, or synthesizeSpeech return, field names, types, or structure. LLMs cannot plan downstream tool calls or extract required data (e.g., voice IDs needed for synthesis) without documented output structure.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 48 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 0 | - | v1 |
synthesizeSpeechLarge is a near-duplicate of synthesizeSpeech, differing only in default model ('lightning-large' vs 'lightning'). This violates single-responsibility and introduces naming confusion. Should merge into synthesizeSpeech with explicit model parameter, or rename one tool to make distinction obvious (e.g., synthesizeSpeechStandard vs synthesizeSpeechLarge).
Error responses from tools do not guide LLM recovery. WavesApiError includes status, reason, and text, but tools likely return raw error strings without actionable next steps. E.g., 'Waves API Error: Status 401' tells agent nothing, should say 'Authentication failed. Verify WAVES_API_KEY environment variable is set and valid.'
deleteClone is destructive but has no confirmation mechanism. No dry-run, no warning in description, no confirmation step. Agents make mistakes, irreversible deletes should require explicit acknowledgment or support a preview before executing.
Tool descriptions are generic and lack intent clarification. E.g., getClones: 'Fetches the list of cloned voices for a specified model.' Does not explain when to call it vs getVoices, what structure is returned, or what voiceId values are suitable for synthesizeSpeech. Descriptions should guide when/why to use each tool.