AI-powered pipeline that converts YouTube videos into polished articles using vision-based transcription
Server exposes 2 tools with STDIO transport only (hard cap 50). Tool definitions are present in source code with schemas and descriptions, but multiple critical issues undermine quality: (1) Input schemas visible and structurally sound for both tools, but descriptions are functional rather than LLM-optimized; (2) The server chains multiple complex operations (download → split → transcribe → write article) but exposes only 2 low-level tools that require orchestration via FastAgent chains rather than simple, composable interfaces; (3) prepare_youtube_video combines download AND normalization AND splitting, violates single-responsibility principle; (4) No error handling guidance for LLMs (e.g., what if video is too large, transcription fails, path is invalid); (5) Return types are documented in description text but not in structured schemas; (6) No output schema formally declared; (7) Security concern: video_path_str is user input with no validation against path traversal; (8) STDIO transport means this server is not remotely accessible and cannot be used with hosted MCP clients, architectural limitation.
Download a YouTube video, normalize the filename, split into parts.
Transcribe a local video file using an internal multimodal agent.
prepare_youtube_video combines three distinct responsibilities (download + normalize + split). Should be split into separate tools (download_youtube_video, normalize_filename, split_video) so agents can compose them as needed.
STDIO transport only. Server is not remotely accessible and cannot be used with hosted MCP clients. Hard architectural limitation, evaluates to max 50 on protocolReadiness.
No output schemas documented. Tool descriptions mention return values (JSON with base_dir and parts, or transcript string) but schema structure is not formally declared. LLMs cannot parse expected response fields.
video_path_str parameter has no path traversal validation. Input string is used directly in Path(video_path_str.strip()).read_bytes(). Malicious input like '../../../etc/passwd' could read arbitrary files.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 32 | - | v1 |
No error handling guidance for LLMs. Error messages are returned (e.g., 'Error: File not found at {video_path}') but do not suggest recovery steps (e.g., 'verify the path, or call list_files to find available videos').
prepare_youtube_video description does not specify what file formats are supported, what the target MB size controls, or what the part_mb default (15 MB) implies for typical videos.
transcribe_video_file description does not document which video codecs/formats are supported, whether transcription is deterministic/idempotent, or expected runtime for different video lengths.
No idempotency guarantees documented. If prepare_youtube_video is called twice with the same URL, does it re-download or return cached parts? Does transcribe_video_file re-process or return cached transcript?
part_mb parameter has a default (15) but no min/max bounds documented. Can agents pass 0, 1000, or 999999? Unbounded numeric parameters allow LLMs to pass absurd values.