podcli has 22 tools with generally adequate naming (verb-based) and descriptions, but several quality gaps limit production readiness. Most tools have input schemas with types and descriptions, but many lack output schema documentation. Parameters are often described at a surface level without constraint details (ranges, patterns, or dependencies). Error handling is minimal, tools don't guide recovery or classify errors. Security concerns: assemblyai_api_key exposed as a parameter (violates secret-injection pattern); file paths accepted directly without validation warnings. Descriptions average ~100 - 150 chars, which is acceptable but could be more action-oriented for agent selection. The largest issue is absent output schemas for nearly all tools, agents cannot plan downstream calls or know what fields will be returned. No tool annotations (readOnly/destructive hints) despite having both read-only (get_ui_state) and destructive (delete_from_knowledge_base) operations.
Tools (22)
add_to_knowledge_basewrite50/100
Add or update a file in the knowledge base
analyze_silenceread onlysource verified68/100
Detect silence and pauses in a video for silence removal optimization
batch_create_clipswritesource verified70/100
Create multiple clips with a bounded worker pool for parallel rendering
create_clipwritesource verified75/100
Create a single finished short-form clip with captions and formatting
delete_from_knowledge_basedestructive50/100
Delete a file from the knowledge base
get_knowledge_baseread only50/100
List or read files from the knowledge base
get_ui_stateread onlysource verified67/100
Read the current UI state including video, transcript, suggestions, and settings
Missing output schemas for all 22 tools. Agents cannot plan downstream calls, extract required fields for chaining, or validate returned data. Without output schema documentation, agents must infer structure from descriptions, which invites misuse and context loss.
assemblyai_api_key exposed as a tool parameter in transcribe_podcast and transcribe_start. API keys must never appear in tool parameters, they end up in agent traces and prompt history. Use server-side secret injection via environment variables or a secrets vault.
Recommendations
Document the return type (JSON schema) for every tool. Example for get_ui_state: {"video_path": string, "transcript": {"text": string, "words": [{"word": string, "start": number, "end": number}]}, "suggestions": SuggestedClip[], "settings": {...}}. This lets agents plan next steps and extract chaining IDs.
Move assemblyai_api_key to server-side configuration (environment variable, .env, or vault). If the user must provide their key, read it from process.env.ASSEMBLYAI_API_KEY at initialization, not as a tool parameter. Document: 'Set ASSEMBLYAI_API_KEY in your environment to enable AssemblyAI transcription.'
Add explicit path validation guidance to file path parameters. Example: 'Absolute path to podcast file. Only paths within the configured media directory are accepted; parent directory traversal (../) is blocked.' Or: 'Relative path from the project root; must not contain '..' or absolute system paths.'
For transcribe_podcast and batch_create_clips, add error recovery guidance in the description: 'May fail if the file is unreadable, the format is unsupported, or the transcription service times out. On failure, check file format and permissions, or try a smaller model_size for faster processing.' This guides agent retry logic.
Add tool annotations to every tool. Example for delete_from_knowledge_base: {"inputSchema": {...}, "annotations": {"destructive": true}}. Mark read-only tools with readOnly: true. This enables agents to reason about side effects.
Clarify list_* and get_* tool purposes with 60 - 120 character descriptions. Example for list_outputs: 'List all rendered clip files in the output directory, with file size and creation timestamp. Call this after batch_create_clips to verify exports succeeded.'
Spec posture evidence
Inferred effective spec: <=2025-11-25.
Relies on Logging (deprecated) - log to stderr or use OpenTelemetry
Score history
Overall score trend
First recorded score · v2 rubric
59/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-23
D
59
<=2025-11-25
v2
Get AI-generated workflow guidance based on current state
import_transcriptread onlysource verified68/100
Import transcript from file (SRT, VTT, or JSON)
job_statusread only50/100
Poll the status of an async transcription job
list_assetsread only50/100
List available logos, intros, outros, and other media assets
list_outputsread onlysource verified55/100
List all rendered clips in the output directory
list_presetsread only50/100
List available clip export presets (caption styles, crop strategies, etc.)
manage_integrationswriteauth50/100
Manage integrations (YouTube, etc.) and configuration
modify_clipwritesource verified68/100
Modify a suggested clip's title, timing, or other properties
parse_transcriptread onlysource verified67/100
Parse raw transcript text into word-level timestamps
render_silence_removedwritesource verified67/100
Create a new video with silence removed based on analysis
set_videowritesource verified62/100
Set the current video file without transcribing
suggest_clipsread onlysource verified68/100
Analyze transcript and suggest viral clip moments
toggle_clipwritesource verified65/100
Select or deselect a clip for batch export
transcribe_podcastwriteauth50/100
Transcribe a podcast episode with speaker detection and word-level timestamps
File paths (file_path, video_path, logo_path, etc.) accepted directly as strings without validation warnings or path traversal guards documented. Agents can be tricked into passing '../../../etc/passwd' or absolute paths outside intended directories. Descriptions should warn about path restrictions or state that only relative paths are accepted.
No error handling guidance. Tools lack descriptions of what can fail, how to recover, or whether errors are retryable. For example, transcribe_podcast might fail if the file is unreadable, the Whisper API times out, or the language is unsupported, but the error message won't tell the agent what to do next.
No tool annotations (readOnly, destructive, idempotent hints). delete_from_knowledge_base is destructive but unmarked. get_ui_state and list_* tools are read-only but unmarked. Without annotations, agents cannot reason about side effects, idempotency, or safe retry strategies.
Vague or missing descriptions for discovery/list tools. list_outputs, list_assets, list_presets, and list_knowledge_base lack clear statements of what they return, when to call them, or what structure they expose. Agents resort to guessing.
Numeric parameters lack ranges. suggest_clips.max_duration defaults to 60 but no min/max bounds documented. create_clip.start_second and end_second have no constraints stated. This invites agents to pass invalid values (negative seconds, durations longer than video).
manage_integrations uses a generic 'action' enum (list, connect, disconnect, get_config, set_config) and an overloaded 'config' object parameter. The description doesn't clarify what config fields are required for each action, or what get_config returns. This forces agents to guess the schema.
Parameter dependency documentation missing. For example, suggest_clips accepts both 'words' (array) and 'duration' (number), but it's unclear if both are required, mutually exclusive, or optional. clip_numbers and clips parameters in batch_create_clips are both arrays, when does the agent use each one?
Decompose manage_integrations into separate tools: list_integrations, connect_integration (action=connect+integration+config), get_integration_config (action=get_config+integration), set_integration_config (action=set_config+integration+config). Or document the config object schema for each integration type.
Document parameter relationships explicitly. For batch_create_clips: 'Either provide clip_numbers (array of indices from suggest_clips results) or clips (array of explicit clip objects with start_second, end_second, title). Passing both is an error.' For suggest_clips: 'If words (word-level timestamps) are provided, they override any parsing of the transcript text; duration is optional if words are present.'
Add pagination to list_outputs, list_assets, list_presets, and get_knowledge_base (list action). Accept limit and offset/cursor parameters; return a total count. Example: list_outputs accepts {limit: 20, offset: 0} and returns {clips: [...], total: 150, has_more: true}.
Create an idempotency pattern for async operations. transcribe_start should accept an optional idempotency_key; if the same key is used within a time window, return the existing job_id instead of creating a duplicate job. Document: 'Pass idempotency_key to avoid duplicate transcription jobs if the network request is retried.'
Add a dry_run or preview parameter to create_clip and batch_create_clips. This lets agents preview the output and confirm before rendering (which may be slow). Example: create_clip(..., dry_run=true) returns {preview_url, estimated_duration_seconds} without rendering.
For suggest_clips, clarify what 'style' values are supported and how they affect clip selection. Document with an enum: {"style": {"type": "enum", "enum": ["viral", "educational", "entertaining", "takeaway"], "description": "Clip selection style. 'viral' prioritizes shocking moments; 'educational' highlights key insights; 'entertaining' finds humorous segments; 'takeaway' extracts actionable advice."}}.