YouTube video/audio/transcript downloader with CLI, Web UI, and MCP support
VidSnatch has 9 tools with moderate quality definitions. Strengths: all tools have descriptions (avg 180 chars, within baseline); most parameters are typed and described; tool names follow verb_noun convention and are action-oriented. Weaknesses: no explicit output schemas documented anywhere (critical gap for schema score); descriptions lack recovery guidance and error context; no parameter constraints (enums, ranges, patterns) for known-set values like quality/format; no documented error categories or actionable failure messages; no tool annotations (readOnlyHint/destructiveHint) despite clear risk separation; descriptions occasionally include procedural workflow instructions instead of focusing on what the tool does. The stitch_videos tool has an inferred definition (appears in mcp_tools.py but implementation code not fully visible), capping it at 50. Tools operating on the same resource (download_video, download_audio, download_video_segment) have clear naming distinctions but lack cross-tool composition guidance in descriptions.
Download audio from a YouTube video to the configured download directory.
Download transcript with timestamps from a YouTube video. This is ESSENTIAL for finding specific topics or segments in videos. The transcript includes precise timestamps for each spoken segment, making it perfect for: - Locating when specific topics are discussed (e.g., "Windsurf deal", "AI features", etc.) - Finding exact time ranges for creating video clips - Searching through long videos to identify relevant sections WORKFLOW TIP: Always download the transcript FIRST when users ask for clips about specific topics, then use the timestamps to determine start_time and end_time for download_video_segment.
Download a YouTube video to the configured download directory.
Download a specific segment/clip from a YouTube video using precise timestamps. IMPORTANT: When users request clips about specific topics (e.g., "download the part about X"), you should FIRST use download_transcript to get the timestamped transcript, then analyze it to find the exact time range when that topic is discussed, and finally use those timestamps here. This tool is perfect for: - Creating short clips from long videos - Extracting specific discussions or segments - Sharing relevant portions without downloading entire videos
No output schemas documented. Tools return JSON strings but no schema definitions specify field names, types, or structure. LLMs cannot infer downstream tool chaining (e.g., what fields does get_video_info return? Can search_videos output be passed directly to get_video_info?).
Quality and format parameters accept free-form strings instead of enums. 'quality' accepts 'highest', 'lowest', or 'specific quality like 720p' without enum constraint, LLMs may hallucinate invalid values like 'best' or '1080p'. Format accepts 'mp3', 'm4a', 'wav' but no enum defined.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) in visible code. Risk ratings are documented (READ_ONLY vs WRITE) but not surfaced as tool annotations in the schema, prevents clients from applying safety rules.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 57 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 43 | - | v1 |
Get the current MCP server configuration.
Get detailed information about a YouTube video including title, duration, and available formats. Use this tool to understand video content before processing. For long videos where users want specific segments, consider following up with download_transcript to get timestamped content that can help locate specific topics or discussions.
List all files in the download directory.
Search YouTube for videos matching a query. Returns up to 10 results. Use this tool to find YouTube videos by keyword before downloading. The returned URLs can be passed directly to get_video_info, download_video, download_audio, or download_transcript.
Stitch multiple video clips into one video. WORKFLOW: Use after download_video_segment() calls. 1. search_videos() → find relevant videos 2. download_transcript() → identify timestamps for subtopics 3. download_video_segment() × N → save each clip 4. stitch_videos(file_paths) → join all clips into one compilation
Error handling not documented. Code catches exceptions and returns JSON with 'error' field, but no guidance on error categories (retryable vs fatal), what the LLM should do next, or valid values for recovery. Example: download_video catches 'Exception' generically, LLM cannot distinguish network timeout (retry) from invalid URL (fail).
Descriptions mix procedural workflow instructions with tool definitions. download_transcript and download_video_segment include multi-step guidance ('FIRST use download_transcript...THEN use...FINALLY use...') instead of focusing solely on what that single tool does. Workflow coordination belongs in the agent prompt, not tool descriptions, which should be concise and self-contained.
get_config has a vague description ('Get the current MCP server configuration') with no indication of what fields are returned or when an LLM would call it. Is it for discovering available download directory, auth requirements, rate limits? Unclear.
stitch_videos tool definition is inferred from code but implementation details not fully visible in provided source excerpt. Parameters declared (file_paths, output_filename) but no schema visibility confirms proper JSON Schema structure.
No pagination or result limits documented for search_videos. Description says 'Returns up to 10 results' but no limit parameter visible in schema, and no mention of how to fetch beyond top 10.
list_downloads lacks context. Returns 'all files in the download directory' but no indication of structure (file list, metadata, sizes?). When would agent use this vs calling download_video/audio directly? Unclear.