YouTube video knowledge engine — transcripts, vision, and persistent wiki
mcptube presents 20 tools with generally clear names and solid descriptions. All tools have explicit descriptions (100-400 chars typically), and all parameters have type definitions and descriptions. However, output schemas are largely undocumented, the codebase shows return types in docstrings but no formal schema documentation that LLMs can parse. Tool names follow verb_noun patterns well (add_video, list_videos, wiki_search). Several tools are passthrough/analysis tools that delegate decision-making to the user ('Returns data for YOU to...'), which reduces the pattern alignment. Error handling is basic: VideoAlreadyExistsError and VideoNotFoundError are caught, but recovery guidance is minimal. Parameter defaults are reasonable (text_only=False, limit=10, page_type=None). Some tools like 'wiki_ask' and 'synthesize' have powerful multi-tool composition potential but lack detailed output schema documentation. The server follows a coherent domain model (Library → Wiki → Frame extraction → Analysis), and tool names clearly distinguish concerns (add_video vs list_videos vs get_info). However, passthrough tools (classify_video, generate_report, ask_video) blur the boundary between tool and UI logic, reducing clarity on what the tool guarantees vs. what the user must do.
Ingest a YouTube video and build wiki knowledge pages.
Returns transcript for YOU to answer about a single video.
Returns transcripts for cross-video Q&A.
Returns metadata for YOU to classify.
Search YouTube for videos on a topic.
Returns data for YOU to write an illustrated report.
Returns multi-video data for cross-video report.
Output schemas are not formally documented. Tools return dicts/lists, but LLMs cannot parse the structure of response fields, they rely on docstring hints like 'Returns: model.model_dump(mode="json")' which is not machine-readable. LLMs cannot plan downstream tool calls without knowing what fields are available.
Passthrough analysis tools (classify_video, generate_report, ask_video, generate_report_from_query, synthesize, ask_videos) blur the boundary between tool and UI. Their descriptions say 'Returns data for YOU to write/answer', they delegate decision-making to the user rather than making a determination. This reduces tool utility and makes it unclear to LLMs what they should do with the results. Either the tool should make the decision (high-value), or the description should clarify that the tool is a data-retrieval helper and the user/LLM must interpret the output.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | D | 55 | 2026-07-28+ | v2 |
Extract a frame from a video at a specific timestamp.
Search a video's transcript and extract a frame at the best matching moment.
Returns base64-encoded frame data for embedding.
Get full details for a video including transcript and chapters.
List all videos in the library.
Remove a video from the library and clean wiki references.
Returns multi-video data for theme synthesis.
Ask a question — answered via agentic wiki retrieval.
View version history for a wiki page.
Browse wiki pages with optional filtering.
Full-text search across all wiki pages.
Read a specific wiki page in full.
Get the wiki table of contents.
Error handling is minimal. Only two error types are caught (VideoAlreadyExistsError, VideoNotFoundError), and they return bare {'error': str(e)}. LLMs cannot distinguish between retryable and permanent failures, or know what recovery steps to attempt. Frame extraction errors are not caught, a FrameExtractionError would crash ungracefully. Missing recovery guidance: 'User not found. Try search_users() first' pattern not applied.
Parameters could accept human-friendly identifiers alongside system IDs. 'video_id' is required as an 11-character YouTube ID, but users typically refer to videos by title or channel. There is no video_name or video_title parameter. This forces extra discovery calls (list_videos → find by title → extract ID). Pattern: natural-identifiers.
Pagination is not offered for list_videos() and wiki_list(). If the library grows large (100+ videos), returning all at once wastes tokens and risks context explosion. wiki_list has optional tag/type filters but no page/limit parameters. list_videos has no parameters at all.
wiki_search() accepts 'limit' but doesn't document if results are ordered by relevance and what the total count is. Without pagination cursors or a total count, the LLM cannot know if there are more results or how to page through them. Also, no mention of whether limit caps at a hardcoded max (e.g., 100) to prevent DoS.