MCP server using Gemini Vision API to interact with YouTube videos.
Server provides 4 tools with basic Zod schemas converted to JSON Schema via zodToJsonSchema. However, critical deficiencies emerge across naming, description quality, and parameter documentation. Tool names lack clear action verbs and descriptions are generic placeholders rather than LLM-optimized guidance. Parameter descriptions are either missing, trivial, or misleading (e.g., youtube_url description is 'Invalid YouTube URL provided.', an error message, not a parameter guide). The 'ask_about_youtube_video' tool conflates two concerns (answering a question vs. providing general description). Output schemas are not documented, LLMs cannot plan downstream calls or extract specific fields. Error handling is present but generic ('Gemini API Error'), without recovery guidance. Security is reasonable (API key injection via env), but tool compositions lack the idempotency and chaining guidance needed for reliable agent use. This is a D-tier server: functional definitions exist, but they do not meet production LLM-agent standards.
Answers a question about the video or provides a general description if no question is asked.
Extracts key moments (timestamps and descriptions) from a given YouTube video.
Lists available Gemini models that support the 'generateContent' method.
Generates a summary of a given YouTube video URL using Gemini Vision API.
Output schemas completely undocumented across all tools. LLMs cannot infer return types, field names, or structure. Downstream tool chaining is impossible without trial-and-error.
Parameter descriptions are misleading or missing. youtube_url description is 'Invalid YouTube URL provided.', an error message, not a parameter guide. This violates pattern:tool-description and confuses LLMs about what input is expected.
ask_about_youtube_video conflates two concerns: answering a specific question vs. generating a general description when no question is provided. This violates pattern:tool (single responsibility). LLM cannot reason about when to use this vs. summarize_youtube_video.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 45 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Tool descriptions lack WHEN/WHY guidance. 'Answers a question about the video' tells LLM the HOW, not the WHEN. Compare to production baseline: LLM-optimized descriptions (50-200 chars) state purpose, prerequisites, and differentiation from similar tools.
Error handling is generic and non-actionable. callGeminiApi() throws 'Gemini API Error: <message>' with heuristic error classification based on message substrings. This provides no recovery guidance (e.g., 'Try again in 60 seconds' for quota errors, or 'Check API key' for auth errors).
No pagination or result limits documented. Tools may return unbounded results (e.g., list_supported_models returns all Gemini models). Large result sets blow context windows and waste tokens. Missing limit parameter and documentation of defaults.
Tool annotations (readOnlyHint, destructiveHint, idempotentHint) are not used. All tools are read-only (per risk classification), but this is not communicated via structured annotations in the tool definition. This omission makes the protocol less current against the 2026-07-28 spec.