An MCP server that transcribes online videos from YouTube and Bilibili with timestamped output. Downloads video audio, converts to WAV format, and uses cloud-based WhisperX models via Replicate for transcription.
Two tools present with verb-noun naming (get_youtube_transcript, get_bilibili_transcript), both reasonably described. However, critical gaps emerge: input schemas are visible and minimal (only 'url' parameter with basic string type and description), but output schemas are completely undocumented, the tools return a 'str' (transcript text) with no structure or field documentation. Parameter descriptions are present but sparse (10-30 chars). Error handling exists in code (ValueError, Exception catches with ctx.error calls) but lacks recovery guidance for LLMs. No input validation constraints (URL format, length limits). Tool names are clear but parameters accept free-form strings with no format enforcement. Descriptions (~100 chars each) meet minimum length but lack WHEN/WHY context for LLM selection. No tool annotations (readOnlyHint present semantically as 'Risk: READ_ONLY' but not declared in schema). Output is unstructured text, violating pattern:response-shaper. This server demonstrates basic structure but falls short of production-grade agent integration.
Provide a bilibili url and obtain a timestamped transcript of that bilibili video
Provide a youtube url and obtain a timestamped transcript of that youtube video
No output schema documented. Tools return unstructured text ('str') with no field definitions. LLMs cannot plan downstream operations or extract structured data for chaining. Violates pattern:response-shaper and pattern:tool requirements.
Input parameter 'url' lacks format constraints. No validation rules (regex pattern, length limits, URL scheme enforcement). Description states 'A url of the youtube/bilibili video' but does not specify format or what constitutes a valid URL. LLMs may pass malformed strings; server must validate and return actionable error messages.
Error handling exists in code (ctx.error calls) but lacks recovery guidance. Error message 'Failed to transcribe video: {str(e)}' tells LLM nothing about next steps. Missing: error classification (retryable vs fatal), suggested alternatives, or guidance on what to do after failure. Violates pattern:recovery-guide.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 43 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 37 | - | v1 |
Tool descriptions lack WHEN/WHY context. Both descriptions state WHAT the tool does ('obtain a timestamped transcript') but do not explain: When to use get_youtube_transcript vs get_bilibili_transcript? What preprocessing happens? What's the expected runtime? What formats are supported? Ambiguity forces LLM to guess selection criteria.
No tool annotations (readOnlyHint, idempotentHint) declared in schema. Risk metadata exists in evaluation ('Risk: READ_ONLY') but is not exposed to MCP protocol. LLM cannot reason about safety without this signal. Add 'readOnlyHint: true' to tool definitions.
Parameter description brevity. 'A url of the youtube video' is only 28 chars. Baseline for A+ tools is 70+ chars per parameter. Should include: format example, supported URL schemes (https://www.youtube.com/...), expected content (public videos only? live streams?), and any platform-specific limitations.