MCP server for Fish Audio Text-to-Speech integration
Fish Audio MCP Server provides two tools with explicitly defined schemas and descriptions. Both tools have clear names starting with action verbs (fish_audio_tts, fish_audio_list_references) and reasonable descriptions (194 and 180 chars respectively). The TTS tool has a comprehensive input schema with 13 parameters, most with type definitions and descriptions. However, there are significant gaps: (1) output schemas are not documented anywhere, the CallToolRequestSchema handler returns generic text responses with JSON stringified results, not structured objects; (2) parameter descriptions lack actionable constraints and format guidance; (3) error handling provides no recovery guidance; (4) no per-parameter type validation is visible in the tool implementations; (5) the responses are wrapped in a generic text container rather than structured typed objects. The ListReferences tool has an empty properties object but no documented output schema. Overall, definitions are above average for community MCP servers but fall short of production-grade tool design.
List all configured voice references
Generate speech from text using Fish Audio TTS API
Output schemas not documented. Tools return generic text responses with JSON stringified results. LLMs cannot plan downstream operations without knowing the response structure. fish_audio_tts should document it returns {success: boolean, audio_data?: string (base64), file_path?: string, format?: string, played?: boolean, streaming_mode?: string, total_bytes?: number, error?: string}. fish_audio_list_references output structure is completely undocumented.
fish_audio_list_references has empty input schema (no properties). While correct for a parameter-less tool, the OUTPUT schema is missing entirely. There is no documentation of what fields the response contains, what structure it has, or how to use the returned references.
Parameter descriptions lack actionable constraints. E.g., 'text' parameter says 'Text to convert to speech' with maxLength 10000, but description does not mention the length limit. 'mp3_bitrate' enum is [64, 128, 192] but description does not explain what these numbers mean (kbps, see enum comment). 'format' enum is [mp3, wav, pcm, opus] but description does not clarify which format is default or when to use each. LLMs cannot reliably select values without explicit constraint descriptions.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 58 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Error handling provides no recovery guidance. The CallToolRequestSchema handler catches errors and logs them, but there is no per-tool error classification or next-step guidance. If API calls fail, the LLM receives no indication of whether to retry, check parameters, ask the user for clarification, or abandon the operation.
Missing parameter descriptions for optional parameters. 'reference_id', 'reference_name', and 'reference_tag' have descriptions, but it is unclear which takes precedence or if passing multiple simultaneously is valid. The relationship between these three parameters is not documented.
Streaming modes (streaming, websocket_streaming, realtime_play) are Boolean flags with minimal explanation. The description for 'streaming' says 'Enable HTTP streaming mode' but does not explain when to use it, performance implications, or whether it is compatible with websocket_streaming.