An MCP server providing high-performance audio transcription using Faster Whisper, supporting batch processing, multiple output formats (VTT, SRT, JSON), and model caching
The server provides 3 tools with documented input schemas and descriptions, but critical gaps prevent higher scoring. Naming uses verb prefixes appropriately (get_, transcribe, batch_transcribe). However, descriptions are minimal (under 100 chars for most), parameter descriptions lack detail about constraints and formats, output schemas are undocumented, and error handling is absent from tool definitions. The codebase delegates to helper modules (transcriber.py, model_manager.py) whose actual implementations are not visible, making it impossible to verify error handling, validation, or output shaping. Parameters like 'device', 'compute_type', and 'output_format' accept specific values but lack enum constraints in the schema. No evidence of input validation, permission checks, or recovery guidance. Tool composition is reasonable (three distinct concerns), but the lack of output documentation and error guidance significantly limits LLM reasoning capacity.
批量转录文件夹中的音频文件
获取可用的Whisper模型信息
使用Faster Whisper转录音频文件
get_model_info_api: Description is only 14 characters ('获取可用的Whisper模型信息'). Does not explain WHAT structure is returned, WHEN to call it, or what fields the LLM should expect. Breaks pattern:tool-description.
All three tools: Output schemas are completely undocumented. Tool definitions specify return type as 'str' only. LLMs cannot plan downstream calls or extract structured data without knowing what fields to expect (e.g., does get_model_info_api return a list of model objects, or a single string summary?). Violates pattern:tool and pattern:response-shaper.
transcribe & batch_transcribe_audio: Parameters with restricted values (device: cpu/cuda/auto; compute_type: float16/int8/auto; output_format: vtt/srt/json) are not declared as enums in the input schema. Schema shows 'type': 'string' only. This invites hallucinated values from LLMs. Violates pattern:constrained-input.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 42 | - | v1 |
transcribe & batch_transcribe_audio: Parameter descriptions lack actionable constraints. E.g., 'beam_size' description says '波束搜索大小,较大的值可能提高准确性但会降低速度' (larger values improve accuracy but reduce speed) but omits minimum/maximum bounds, default rationale, or recommended range. Parameter descriptions average 72 chars in A+ tools; many here are 30-50 chars without specifics. Violates pattern:constrained-input and review:param-validation-rules.
Error handling entirely absent from tool definitions. No guidance on what happens on file not found, unsupported format, GPU out of memory, invalid model name, or transcription failure. Violates pattern:recovery-guide and pattern:error-classification.
transcribe: 'audio_path' parameter accepts free-form string. No validation guidance. Description does not specify accepted formats (audio_processor.py shows support for mp3, wav, m4a, flac, ogg, aac, but this constraint is invisible to LLM). LLM cannot know which formats are valid without calling the tool and failing. Violates review:param-validation-rules.
batch_transcribe_audio: 'parallel_files' description says '仅在CPU模式下有效' (only valid in CPU mode) but the parameter has no conditional logic expressed in schema. If user passes device='cuda' and parallel_files=4, will it be silently ignored or error? Undocumented dependency. Violates review:param-relationships.
Helper modules (transcriber.py, model_manager.py, formatters.py) are referenced but their code is not shown. Cannot verify if input validation, error handling, or recovery logic actually exist at runtime. Only audio_processor.py and formatters.py snippets are visible, showing validation and formatting but no error recovery or actionable error messages.