Single-tool server with significant definition quality gaps. The tool 'generate_music_suno' has a verbose description (372 chars, within baseline range), but the schema lacks required parameter declarations and parameter descriptions are inconsistent. The schema properties are defined with types and some descriptions, but the 'required' array is empty despite conditional logic indicating prompt/tags/title OR gpt_description_prompt should be required. The tool includes HTML/audio formatting instructions in the description (anti-pattern: instructions belong in implementation, not description). No output schema is documented. Error handling exists but lacks recovery guidance. The server is STDIO-only, which is a hard constraint capping protocol readiness at 50, but definition quality issues stand independently.
Generates a song using the Suno API. Provide lyrics, style, and title for custom mode, or a description for inspiration mode. Returns the audio URL upon completion. Polling for results may take a few minutes. When returning an audio URL, please use the following HTML format for user convenience: ```html <audio controls> <source src="YOUR_AUDIO_URL_HERE" type="audio/mpeg"> </audio> <br> <a href="YOUR_AUDIO_URL_HERE" download="SONG_TITLE.mp3"> 点击这里下载喵! </a> ```
No required parameters declared in schema. The tool accepts both custom mode (prompt, tags, title) and inspiration mode (gpt_description_prompt), but the schema has empty required array. This forces the LLM to guess which parameters are mandatory. Validation happens at runtime but schema does not express the constraint.
No output schema documented. The tool description mentions 'Returns the audio URL upon completion' and includes HTML formatting instructions, but there is no formal schema showing what fields are returned (e.g., does it return {url, task_id, status}? {audio_url}? Plain string?). LLMs cannot plan downstream calls or extract structured data without knowing the response shape.
Tool description includes implementation details (HTML audio tag code, Chinese text '点击这里下载喵!'). These instructions should not be in the tool description, they pollute the LLM prompt and should be handled by the client rendering layer or agent instructions. The description should state WHAT the tool does and WHEN to use it, not HOW to render the output.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 48 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Conditional parameter dependencies not documented in schema or parameter descriptions. The description mentions 'If provided, continue_at and continue_clip_id are also required' for task_id, and 'If provided, prompt, tags, and title are not strictly required' for gpt_description_prompt. These dependencies are buried in parameter descriptions rather than expressed in a clear, machine-readable way. LLMs must parse natural language to infer relationships.
Parameter descriptions include example values (e.g., '[Verse 1]\nUnder the starry sky...' for prompt, 'acoustic, folk, pop' for tags). This is an anti-pattern, LLMs frequently reuse example values literally rather than adapting them to the actual context. Constraints (string maxLength, pattern) should be used instead of textual examples.
Polling behavior mentioned in description ('Polling for results may take a few minutes') but no polling parameters (max_wait, timeout, poll_interval) exposed to the LLM. The agent cannot control how long to wait or how often to check. This forces synchronous blocking behavior, which is poor for agent planning.
No error recovery guidance in error responses. The code throws McpError with messages like 'Input parameters are invalid nya~! Please check the requirements...' but does not tell the LLM which parameter is invalid or how to fix it. A detailed error response should say 'Invalid parameters: missing one of [prompt+tags+title] or [gpt_description_prompt]'.