MCP Server for analyzing Instagram videos using Google Generative AI (Gemini), providing video analysis, transcription, and question-answering capabilities
This MCP server has inconsistent tool definition quality with several critical gaps. While some tools (analyze_video_from_url, transcribe_video) have well-structured schemas and reasonable descriptions, others lack sufficient parameter documentation or unclear naming conventions. The server defines 11 tools across two files (Video-Understander/src/mcp_server.py and mcp-server/mcp_sse_server.py), but code inspection reveals incomplete schema enforcement and missing error handling patterns. Tool descriptions range from adequate (194 char baseline met by most) to minimal. Critical issues: (1) Parameters like 'analysis_type' and 'summary_type' lack sufficient context about when to use each enum value; (2) Several tools (get_job_status, get_system_stats) have minimal descriptions that don't explain WHEN an LLM should call them; (3) No visible output schema documentation for most tools; (4) Error handling is absent from visible definitions.
Analyze an Instagram video using AI
Analyze local video file
Download and analyze video from URL (YouTube, Instagram, TikTok, etc.)
Ask specific question about video content
Cancel a job
Extract specific scenes or moments from video
Get the status of an analysis job
Missing or minimal tool descriptions violate pattern:tool-description. Tools like get_system_stats (23 chars), cancel_job (13 chars), and get_job_status (31 chars) fail to explain WHEN to call them or WHAT the output means. LLMs cannot reason about tool selection without context.
Output schemas are not documented for any tool. LLMs cannot plan multi-step workflows without knowing what fields to expect. Missing return type documentation violates pattern:tool baseline ('100% of A+ tools have documented return types').
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 62 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Get system statistics
List recent video analyses
Generate video summary
Get video transcription with timestamps
Enum parameter values lack descriptions. Tools like summarize_video (brief|detailed|bullet_points|chapters) and extract_scenes don't explain the semantic difference between options. LLMs must guess when to use 'brief' vs 'detailed' or 'bullet_points' vs 'chapters'.
Free-form string parameters like scene_criteria (extract_scenes) and custom_prompt (analyze_video_from_url) invite hallucinated input. No constraints, examples, or guidance on expected format. LLMs will try arbitrary prompts that may fail silently.
Error handling is absent. No visible error categorization, recovery guidance, or actionable error messages. If a tool fails (invalid URL, unsupported video format, job not found), the LLM has no recovery path.
Destructive tool (cancel_job) lacks confirmation or dry-run support. An LLM could cancel a long-running analysis without understanding consequences. No safety gate or confirmation step.
Inconsistent enum definitions across similar tools. analyze_instagram_video has analysis_type enum [comprehensive|summary|transcription|visual_description] but analyze_video_from_url has [comprehensive|transcription|summary|visual_description|question_answering]. LLM may pick wrong tool or assume feature doesn't exist.
No pagination guidance for list_recent_analyses. Limit parameter has no min/max bounds. If an LLM requests limit=10000, server may crash or timeout. No total_count or next_cursor returned (not visible in definition).
Parameter type hints are missing or ambiguous. 'language' param in transcribe_video doesn't specify format (ISO 639-1 code, full name, etc.). LLM will guess and may pass invalid values.