Text-to-speech with voice cloning. Supports multiple models (standard, turbo, fish) with emotional markers and paralinguistic tags.
SolSpeak TTS demonstrates good overall definition quality with properly registered tools, comprehensive parameter schemas, and detailed descriptions. All 6 tools have documented input schemas with type constraints and descriptions. However, there are notable gaps in error handling guidance, output schema documentation, and some parameter descriptions lack explicit constraints. Tool naming follows verb_noun conventions (text_to_speech, list_voices, delete_voice, clone_voice_from_youtube, set_voice_transcript, generate_conversation). The server excels at parameter documentation with enum constraints (models: standard|turbo|fish) and numeric ranges (exaggeration, cfg_weight, temperature bounds all specified). Descriptions average ~120 chars per tool, which is adequate but could be more LLM-optimized. The MCP_INSTRUCTIONS provide excellent contextual guidance for multi-model usage and emotion markers for Fish model, but this pedagogical content should be supplemented with inline parameter descriptions for error recovery patterns.
Clone a voice directly from a YouTube video. Include transcript for fish model support.
Delete a saved voice from the voices directory. Requires password.
Generate a multi-voice conversation as a single audio file.
List all saved voices available for voice cloning.
Set or update the transcript for a saved voice. Required for fish model voice cloning.
Generate speech from text using SolSpeak TTS. Returns download URL.
Missing error handling and recovery guidance. Tools lack 'what to do if this fails' instructions. E.g., text_to_speech could fail if voice_name doesn't exist or model is incompatible, but no error classification or retry strategy is documented.
Output schema not documented. None of the 6 tools explicitly declare their return structure. text_to_speech 'Returns download URL' but the JSON structure (URL field name, content-type, format) is not specified. Agents cannot reliably extract or compose subsequent calls without knowing response field names.
generate_conversation 'items' parameter lacks explicit schema for array element structure. Description says 'Each: {text, voice_name, model, exaggeration, cfg_weight, voice_text, temperature}' but this should be a formal nested schema with type=object, required fields, and per-field descriptions. Current format forces LLMs to infer structure from prose.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 62 | 2026-07-28+ | v2 |
| 2026-03-09 | C | 62 | - | v1 |
delete_voice requires password parameter, security risk. Passwords should never be tool parameters; they belong in server-side secret injection or authentication headers. Storing passwords in agent traces and logs violates the pattern:secret-injection. This tool should use permission-gate pattern instead.
Conditional parameter dependencies not documented. E.g., voice_text is 'REQUIRED for fish model voice cloning' but schema shows nullable:true with no conditional enforcement. voice_audio_base64 conflicts with voice_name (cloning from file vs. saved voice) but mutual exclusivity is not declared. LLMs will pass both or neither, causing silent failures.
No idempotency guarantees documented. Agents retry on ambiguous failures, if text_to_speech is called twice with same input, does it return the same URL or create duplicate audio files? This affects agent reliability and cost.