Static source inference · medium confidence · detected: Logging
Deprecated protocol patterns detected
Summary
The server defines 4 text-to-speech tools with reasonable parameter coverage and basic descriptions. However, there are significant gaps: (1) descriptions are generic and lack LLM guidance on when/why to choose each tool; (2) parameter descriptions are present but minimal (10-50 chars); (3) output schemas are documented informally in code but not explicitly declared; (4) error handling is minimal (catch-all returns 'status: error' with raw exception string); (5) no enum constraints on 'voice' or 'profile' parameters despite being preset/lookup values; (6) no guidance on mutual exclusivity of the three voice-selection tools (text_to_speech vs text_to_speech_with_voice vs text_to_speech_with_profile). Tool naming is acceptable (verb_noun pattern) and schemas are visible in code, but descriptions need 2-3x expansion and parameter handling needs formalization.
Tools (4)
text_to_speechwritesource verified61/100
Generate speech from text using smart voice selection.
Descriptions are too short (33 - 65 chars) and lack LLM-optimized guidance. Each tool description should answer: What does it do? When/why choose this over related tools? What are the prerequisites? Current descriptions are generic and force LLMs to guess intent.
The 'voice' and 'profile' parameters are free-form strings with no enum constraint, despite representing a known set of preset names. This invites hallucinated voice names. Parameter descriptions include example values ('belinda', 'male_en_british') which LLMs may treat as the only valid options.
No discovery tool (e.g., list_available_voices or get_voice_profiles) to help LLMs learn valid voice/profile names. The codebase loads profiles from 'profile.yaml' but this is not exposed as a tool. LLMs cannot enumerate options without a discovery mechanism.
Recommendations
Expand all tool descriptions to 100 - 200 chars. Include: (a) What the tool does, (b) When to use it instead of similar tools, (c) Prerequisites or dependencies. Example: 'Generate speech from text with automatic voice selection based on emotional tone. Use this for simple single-speaker audio when you don't have a specific voice in mind. Requires text input only; outputs WAV file.'
Add a list_available_voices or get_voice_profiles tool that returns available preset voice names, descriptions, and example audio clips. Expose the profile.yaml data as a queryable resource.
Convert 'voice' and 'profile' parameters to enum types. Query the profiles.yaml file at server initialization and populate the enum dynamically. If enums are not supported, use a string with a regex pattern constraint and add validation logic to return actionable errors like 'Invalid voice: got "john_deere", did you mean "john_smith" or "jane_doe"?'
Clarify the distinction between the three voice-selection tools in their descriptions: (a) text_to_speech = automatic voice selection based on text content; (b) text_to_speech_with_voice = clone a specific preset voice by name; (c) text_to_speech_with_profile = describe desired voice properties in natural language. Add a note in each description saying 'Mutually exclusive with [other tools].'
Implement structured error handling with error categories: (a) Retryable errors (e.g., model loading timeout) → include retry_after guidance; (b) User-fixable errors (e.g., invalid voice name) → include 'did you mean' suggestions and list valid options; (c) Fatal errors (e.g., disk full) → explain the problem and suggest contacting support. Example: '{"status": "error", "code": "invalid_voice", "message": "Voice not found: got \"belinda\". Available voices: [list]. Did you mean \"belinda_v2\"?"}'
Spec posture evidence
Inferred effective spec: <=2025-11-25.
Relies on Logging (deprecated) - log to stderr or use OpenTelemetry
No explanation of when to use text_to_speech (smart voice selection) vs text_to_speech_with_voice (preset voice) vs text_to_speech_with_profile (natural language description). The three tools overlap in purpose but descriptions do not clarify the distinction. LLMs will waste reasoning cycles choosing between them.
Error handling is minimal: catch-all exception handler returns {'status': 'error', 'error': str(e)}. This gives the LLM no recovery guidance. Errors should be categorized (retryable, user-fixable, fatal) and include suggestions for next steps.
Output schema is not formally declared. The tools return dicts (status, audio_path, sampling_rate, duration, text) but this structure is implicit in the code, not documented in tool metadata. LLMs cannot reliably extract fields from an undocumented response.
Parameter descriptions are minimal (10 - 50 chars) and lack actionable constraints. E.g., 'Top-p sampling parameter' tells the LLM nothing about what values are sensible in practice. Descriptions should include range guidance ('0.0 - 1.0 controls diversity; 0.95 is typical for speech') and use cases.
'scene_prompt' parameter has default value 'Audio is recorded from a quiet room.' but accepts arbitrary strings with no validation. No enum or pattern constraint. Description does not explain valid scene types or expected format.
Document output schema formally. Add to tool metadata or README: 'Returns object with fields: status (success|error), audio_path (string, path to generated WAV), sampling_rate (int, Hz), duration (float, seconds), text (string, input text used). On error: status=error, error=string with recovery guidance.'
Expand parameter descriptions with practical guidance. Example for 'temperature': 'Sampling temperature (0.0 - 2.0). Lower values (0.3 - 0.5) produce consistent, predictable speech. Higher values (1.0 - 2.0) increase variability and emotion. Default 0.3 is recommended for most use cases.'
Add validation logic for 'scene_prompt'. Either: (a) Accept an enum of predefined scenes (quiet_room, noisy_office, outdoor, etc.), or (b) If free-form, validate length and warn on unusual inputs. Return clear error: 'Scene prompt must be 10 - 500 chars. Got: "[value]".'
Add a note in the multispeaker tool description explaining the [SPEAKER*] tag syntax with a concrete example: 'Input: "[SPEAKER0] Hello! [SPEAKER1] Hi there! [SPEAKER0] How are you?" Supports up to 10 speakers (SPEAKER0 - SPEAKER9). Voice assignment: if voices param provided, map speakers to voices in order; otherwise use automatic selection.'
Consider batch variants: if agents frequently call the same tool multiple times in sequence, offer a batch version (e.g., text_to_speech_batch accepting an array of texts and returning multiple audio files). This reduces token overhead and latency.