TubeMCP defines 2 tools with clear verb-based names (youtube_search, youtube_get_transcript) and reasonable descriptions (190+ chars each). Both tools have input schemas with typed parameters and descriptions. However, there are significant gaps: (1) output schemas are not explicitly documented in the code, return types are inferred from runtime behavior (list[dict] | str, dict | str) rather than declared; (2) error handling returns unstructured strings instead of actionable error guidance; (3) no pagination support despite search returning potentially large result sets; (4) missing parameter constraints (e.g., max_results_per_query lacks min/max bounds); (5) tool descriptions lack dependency hints and are somewhat generic; (6) no tool annotations (readOnlyHint, idempotentHint). The server avoids multi-concern tools and uses natural identifiers (video_url accepts URL or ID), which is good.
Fetch the transcript and metadata for a single YouTube video. Accepts a URL or video ID. Results are cached locally. Fetch only the videos most relevant to the user's question — avoid bulk fetching. The transcript field is an array of segments, each with text, start (seconds), and duration (seconds).
Search YouTube with multiple queries for broader coverage. Each query runs a separate search; results are deduplicated by video ID. Use 2-3 queries from different angles for best results (e.g. a specific query, a broader one, and an alternative phrasing). Returns metadata only — no transcripts.
Output schemas not documented. Both tools return dict | str (on error) but no formal schema is declared. LLMs cannot plan downstream operations without knowing the structure of returned fields (transcript array, metadata fields, etc.).
Error handling returns plain strings ('Invalid YouTube URL or video ID: {exc}', 'TubeMCPError str(exc)') instead of structured error objects with recovery guidance. LLMs cannot distinguish between retryable and unrecoverable failures or take corrective action.
max_results_per_query parameter lacks bounds or constraints. Default is 3, but no min/max is declared. LLMs could pass unbounded values, risking API abuse or context window overflow.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 55 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 48 | - | v1 |
No pagination support on youtube_search despite potentially returning many results. Results are deduplicated but capped at max_results_per_query per search. Tool description mentions '2-3 queries' but provides no limit or next_cursor mechanism for large result sets.
Tool descriptions lack dependency hints. youtube_get_transcript says 'Fetch only the videos most relevant' but does not explain: 'Call youtube_search first to discover videos.' This forces agents to infer the workflow.
No tool annotations. Both tools are read-only and idempotent, but tool definitions lack readOnlyHint and idempotentHint attributes. LLMs may treat them as potentially destructive.