A MCP server to transcribe audio files using OpenAI Api
The server provides a single tool 'transcribe_audio' with a properly structured input schema and basic description. However, the tool lacks output schema documentation, error handling guidance, and parameter descriptions are incomplete. The description is functional but minimal (59 chars), below the baseline of 194 chars. No tool annotations present. The schema itself is present and typed, but parameter descriptions lack depth regarding constraints, format requirements, and error recovery paths. File I/O operations expose path traversal risks that should be documented in security notes.
Transcribe an audio file using OpenAI Whisper API
Missing output schema documentation. Tool returns transcribed text as 'text' field but does not document the expected response structure, field types, or what the LLM should expect.
Tool description is only 59 characters, well below the baseline of 194 chars (p10=34, p90=392). Missing: WHEN to use this tool, what file formats are supported, what prerequisites exist (OpenAI API key), and failure modes.
Parameter 'filepath' description lacks critical format/constraint details. Should specify: absolute vs relative paths, supported audio formats (mp3, wav, m4a, etc.), maximum file size limits, and path encoding requirements.
Parameter 'language' description mentions ISO-639-1 format (e.g. 'en', 'es') as examples but does not declare an enum of supported languages. LLMs will hallucinate invalid language codes. Should enumerate supported values or link to OpenAI Whisper's language list.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 56 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 38 | - | v1 |
No error handling guidance in tool description. Code catches file-not-found and permission errors, but returns them as plain text in 'isError: true' responses. LLM cannot distinguish retryable errors (temp unavailability) from user-fixable errors (wrong path) from fatal errors (invalid API key).
Path traversal vulnerability: code decodes URIs and strips backslashes ('filepath.replace(/\\/g, '')') but does not validate against directory traversal attacks (../../../etc/passwd). Should use path.resolve() and verify resolved path is within allowed directories.
Missing tool annotations. The tool performs a write operation (save_to_file parameter) and makes external API calls. Should declare 'destructiveHint: true' for the save_to_file behavior and 'idempotentHint: false' to warn that repeated calls to the same audio file will produce identical transcriptions but will overwrite saved .txt files.
No rate limiting documented. Tool calls external OpenAI API which has per-minute token limits. No guidance provided on how agents should handle rate limit errors (429 responses).