A Model Context Protocol server for FFmpeg media processing
The server defines 8 tools with reasonable naming conventions (verb_noun pattern) and adequate parameter schemas. However, descriptions are sparse and lack LLM-optimized depth. Tool descriptions average ~80 chars (well below the 194-char baseline for A+ tools). Parameter descriptions are present but minimal. Output schemas are not documented. No error handling guidance is provided. Most critically, the server lacks composition patterns and does not guide the LLM on when/why to use each tool vs. alternatives (e.g., trim vs. change_resolution both modify videos). The caching/force_run pattern across all mutation tools is non-standard and underdocumented.
Retrieve the final output of a completed job. This tool should only be called after the job status is 'success'. If the job is still 'queued' or 'running', this will return the current status. If the job has 'failed', it will return the error details.
Get the latest status for a job_id.
Enqueue a EXTRACT_AUDIO job. Extract the audio from a given video and in the format requested by user. Returns immediately with job_id + status.
Enqueue a CHANGE_FORMAT job. Converts video to another container (e.g. mp4, mkv, avi) without re-encoding. Returns immediately with job_id + status.
Enqueue a CHANGE_RESOLUTION job. Converts the video resolution according to the user's provided height and width. Returns immediately with job_id + status.
Enqueue a CHANGE_SUBTITLE_FORMAT job. Converts the subtitle format according to the user's provided details. Returns immediately with job_id + status.
Tool descriptions are too brief and lack LLM-optimization context. Most descriptions are 30-50 chars (well below 194-char baseline). They state WHAT the tool does but not WHEN to use it vs. similar tools or what prerequisites exist.
No output schemas are documented for any tool. The rubric requires 'Document the output schema. LLMs need to know what fields to expect so they can plan downstream tool calls.' Job mutation tools return job_id and status, these should be explicitly typed in the tool definition.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 43 | - | v1 |
Enqueue a TRIM job. Returns immediately with job_id + status.
Enqueue a transcript job using Whisper on a video or audio file. Returns immediately with job_id + status.
The force_run parameter is non-standard and poorly documented. Its purpose (cache-bypass) is inferred, not explicit. Parameter description says 'if we need to force run a job or not ignoring cache', awkward phrasing that does not explain cache semantics or why an LLM would use it.
Tool composition is poor. The LLM cannot tell which tool to use for different video operations. There is no discovery tool (e.g., list_available_formats, describe_video_capabilities). An LLM must guess whether change_resolution or trim is better for a given use case, the descriptions do not guide this choice.
No error handling guidance is provided. Tools may fail (file not found, invalid format, ffmpeg crash) but there is no recovery pattern documented. Error responses are not visible in the tool definitions, LLMs will not know what to do if a job fails.
Parameter descriptions are inconsistent and sometimes incomplete. E.g., target_format in start_subtitle_format_change has no enum or list of valid formats. start_audio_extraction defaults target_format to 'mp3' but does not list valid audio formats. This forces LLMs to guess at valid input values.
The start_video_transcription tool accepts 6 parameters. Per-baseline, avg params is 4 (p90=8). While within range, the model/language/output_format parameters could be enums with restricted values to prevent hallucinated inputs (e.g., model='invalid_model').