Enhanced YouTube Transcriber MCP Server — dual-engine transcription (captions + ElevenLabs Scribe), speaker diarization, persistence.
The server implements 7 tools with mostly complete schemas and reasonable descriptions. Naming follows verb_noun conventions (transcribe_, add_, search_, list_, get_, remove_, clear_). Schemas are present and typed for all tools. However, descriptions are sometimes generic, parameter annotations lack sufficient context for LLM reasoning, and error guidance is minimal. The server lacks output schema documentation, tool annotations (destructiveHint/readOnlyHint), and recovery guidance in error cases. Tool composition is reasonable but could be tighter, clear_all_videos and remove_video are destructive operations without confirmation patterns.
Add a YouTube video by URL for transcription and storage.
Clear all loaded videos from memory and storage.
Retrieve the full transcript of a loaded video by ID.
List all currently loaded videos with metadata.
Remove a video from the loaded videos and storage.
Search for a query across loaded video transcripts with context.
Transcribe a YouTube video from a natural language command containing a URL.
Destructive operations (remove_video, clear_all_videos) lack confirmation patterns and dry-run support. No explicit warnings in descriptions that these are irreversible.
Error responses are text-only and lack actionable recovery guidance. Errors return plain strings with no structured error classification (retryable vs fatal vs user-fixable).
No output schemas documented. LLMs cannot plan downstream operations or extract specific fields from responses. For example, list_videos returns a summary but structure is not formally declared.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 66 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 28 | - | v1 |
Tool descriptions lack clear WHEN to use them vs similar tools, and missing dependency hints. E.g., 'transcribe_youtube' description doesn't explain it wraps 'add_youtube_video' or when to use natural language command vs direct URL.
Parameter descriptions lack format constraints and ranges. E.g., 'contextWords' accepts 0 - 20 but description only says 'default: 5, max: 20' without explaining what happens at 0 or why max is 20.
Tool annotations (readOnlyHint, destructiveHint, idempotentHint) are missing. Clients cannot infer which operations are safe to retry or call in parallel.
list_videos has no pagination support. Description states 'List all currently loaded videos' but does not document expected result size limits or pagination strategy if many videos are loaded.
get_video_transcript accepts format enum but description is sparse. Does not explain when to use 'text' vs 'json' or what the JSON structure contains.