An MCP server that downloads audio from YouTube videos and transcribes them using OpenAI's Whisper model.
Single tool 'transcribe_audio' has fundamental definition quality issues. Tool name is acceptable but description is in French and lacks LLM-critical details. Input schema present but parameter descriptions are minimal (only 24 chars). Output schema is entirely undocumented, the tool returns a dict with 'success', 'transcript', or error fields, but this structure is never formally declared. Error handling exists but returns verbose dicts that don't match standard structured output. The server demonstrates basic functionality but fails to meet production-grade definition standards.
Télécharge l'audio d'une vidéo YouTube et le transcrit.
Description is in French ('Télécharge l'audio d'une vidéo YouTube et le transcrit'), unintelligible to English-speaking LLMs and MCP client implementations. Must be in English.
Output schema is completely undocumented. Tool returns dict with keys 'success', 'transcript', 'error', 'error_type', 'human_readable', but LLM has no way to know this structure. Must document return type in tool description or metadata.
Input parameter 'url' has only 24-character description ('URL de la vidéo YouTube'). Must include format constraints, example format, and expected behavior for invalid URLs.
Tool description (63 chars in French) does not answer key questions: When should the LLM call this vs other transcription tools? What formats are supported? What's the typical latency? What are the failure modes? Descriptions should be 10 - 1024 chars and explicitly state WHAT, WHEN, and WHAT IT RETURNS.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 34 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 38 | - | v1 |
API key authentication is enforced via @require_api_key decorator but the mechanism (CLIENT_MCP_API_KEY env var) is not documented in the tool description. LLMs cannot know this tool requires authentication to function.
Error responses return verbose dicts with fields like 'error_type' and 'human_readable', but tool description does not document this structure. Error responses must be actionable ('Try again with a valid YouTube URL') not verbose metadata.
No timeout or resource limit constraints documented. Tool downloads and transcribes audio, this can take minutes and consume significant CPU/disk. Description should state typical latency and any constraints (max video duration, file size limits).