YouTube MCP has 4 tools with solid verb-starting names and reasonable descriptions, but significant gaps in parameter documentation and error handling guidance. All tools are properly registered in FastMCP with input schemas visible in main.py. Descriptions range from 160-350 chars (within the 10-1024 baseline), but parameter-level documentation is inconsistent. Error handling raises ToolError but provides minimal recovery guidance. Output schemas are documented in docstrings but not formally structured in the response. The server follows a coherent design pattern (transcript → analysis) but lacks some production-grade safeguards around secret injection and API key exposure.
Retrieves the transcript/subtitles for a specified YouTube video. Uses the YouTube Transcript API to fetch time-coded transcripts in the requested languages. The transcripts include timing information, text content, and duration for each segment.
Answers natural language questions about a YouTube video's content. This tool leverages Google's Gemini 2.0 Flash model to provide responses to questions based solely on the video's transcript. It extracts insights, facts, and context without watching the video itself.
Searches YouTube for videos matching a specific query and returns detailed metadata. This tool performs a two-step API process: 1. First searches for videos matching the query 2. Then fetches detailed metadata for each result (title, channel, views, etc.)
Generates a concise summary of a YouTube video's content using Gemini AI. This tool first retrieves the video's transcript, then uses Google's Gemini 2.0 Flash model to create a structured summary of the key points discussed in the video.
API keys exposed in environment setup; GEMINI_API_KEY and YOUTUBE_API_KEY appear as plaintext env vars in code. No per-tool scope declaration (read:transcript vs write:search not distinguished). Credentials should be injected server-side and never logged.
Error handling is minimal. ToolError exceptions catch all failures and return string messages without actionable recovery guidance. E.g., 'Transcript error: {str(e)}' tells LLM nothing about retryability, prerequisites, or alternatives.
Output schemas documented only in docstrings, not formally declared. LLMs cannot parse return type structure from plain text. Returns use generic List[Dict] with 'type' and 'data' keys, but structure inside 'data' is not schema-validated.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 57 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 16 | - | v1 |
Parameter descriptions are sparse. 'video_id' described as '11-character string from video URL' is helpful, but 'query' in youtube/query has only 'Natural language question about the video content', no constraints on length, format, or complexity guidelines for LLM.
Hardcoded fake API key in main.py: FAKE_PPLX_API_KEY='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' appears unused but is a credential leak risk if ever deployed or exposed in version control.
youtube/search 'max_results' parameter defaults to 5, capped at 50, but no explanation of why or what happens if LLM requests >50. No guidance on pagination or iteration for large result sets.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint). All four tools are read-only, but this is undeclared, LLMs cannot infer safety/retry semantics without explicit hints.