Video Frame Extraction & Composition Service for AI Vision Pipelines
Strong foundation with clear naming (verb_noun pattern), comprehensive parameter schemas using Zod, and good descriptions. Tool names are action-oriented (capture_*, list_*, download_*). All 7 tools have descriptions (avg ~120 chars) and input schemas with type constraints. However, output schemas are not formally documented in the code, responses are JSON-stringified but lack explicit schema declarations. Error handling returns actionable messages via errorResult(). Tool composition is sound: extract → compose workflow is well-designed. Minor gaps: some parameter descriptions could be more prescriptive about constraints (e.g., fps range 1-30 is in schema but not all descriptions emphasize it); no explicit idempotency guarantees stated.
Compose extracted frames into grid sheets. Provide either job_id (from a previous extraction) or frames (array of file paths), not both.
Discover HLS stream URLs from a webpage or media URL.
Download a video from a URL with optional format selection and merging.
Extract frames from a video at a specified FPS rate. Downloads URLs via yt-dlp automatically. Returns absolute file paths to extracted frames.
Extract frames from a video and compose them into grid sheets in one operation. Downloads URLs via yt-dlp automatically. Returns both frame and sheet file paths.
List available video formats and quality options for a given URL using yt-dlp.
Get video metadata (duration, resolution, FPS, codec) from a local file or URL. ffprobe handles URLs natively — no download needed.
Output schemas not formally documented. Tools return JSON via textResult() but no explicit schema declarations visible in tool registration. LLMs cannot reliably parse response structure without documented schemas.
capture_extract_and_compose combines two operations (extract + compose) in one tool. While convenient, this violates single-responsibility principle and prevents agents from composing these steps independently.
Parameter descriptions for capture_extract and capture_download_video lack explicit guidance on preset/format interactions. E.g., 'width' and 'quality' only apply when preset=custom, but this conditional dependency is not clearly stated in parameter descriptions.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | B | 79 | <=2025-11-25 | v2 |
No explicit idempotency guarantees documented. Tools like capture_download_video and capture_extract may have side effects (file creation, temp cleanup). Agents need to know if retrying with identical inputs is safe.
Error messages are generic. acquireSource() returns errorResult(err) which may expose stack traces or internal details. Error responses should be user-friendly and actionable (e.g., 'Video duration exceeds 3600s. Try a shorter clip or lower FPS.').