Universal API wrapper that automatically wraps any Python library with a robust REST API. Includes MCP server implementation for MLX Whisper transcription and support for YOLO, LLMs, Streamlit, OpenCV, Pandas, and other Python libraries.
This MCP server exposes 3 tools from MLX Whisper. All three tools have schemas and descriptions present. However, definitions have significant gaps: (1) Tool descriptions are terse and lack actionable context (e.g., 'Transcribe audio data to text using Whisper' is 43 chars, below the 50-char practical minimum for LLM discernment). (2) The transcribe_audio tool has a complex oneOf schema with two branches (audio_path vs audio_base64), but the schema lacks clarity on when to use each; descriptions do not explain the trade-off or provide guidance. (3) Parameters lack sufficient descriptions, e.g., 'language' has only 'Language code (e.g., 'en')' which violates the pattern of describing constraints and defaults separately from examples. (4) Error handling is not documented, no guidance on what happens if audio is corrupted, language is unsupported, or the model fails to load. (5) Output schemas are not visible in the provided source, we can infer transcribe_audio returns text, but the exact structure is unknown. (6) No mention of timeouts, rate limits, file size constraints, or supported audio formats. The tools are functional but fall short of production-grade documentation.
Check Whisper API health status
List available Whisper models
Transcribe audio data to text using Whisper
Tool descriptions are too terse (<50 chars) to provide sufficient context for LLM tool selection. E.g., 'Transcribe audio data to text using Whisper' (43 chars) and 'List available Whisper models' (30 chars) lack detail on when to call them, what they return, and what prerequisites exist.
No output schemas documented. Callers cannot see what fields transcribe_audio returns (just text? with metadata?), what list_whisper_models returns (array of strings? objects with size/params?), or what check_whisper_health returns (status code? latency? model load times?). Output schema is critical for LLM planning.
Parameter descriptions are minimal and mix examples with constraints. E.g., 'Language code (e.g., 'en')' does not explain valid format, which languages are supported, or what happens if an unsupported language is passed. Should separate format/constraint documentation from examples.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 48 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 15 | - | v1 |
transcribe_audio uses oneOf with two mutually exclusive branches (audio_path vs audio_base64) but lacks clear guidance on when each is appropriate. No mention of file size limits, supported formats (WAV, MP3, FLAC?), or maximum audio duration. LLMs will guess.
No error handling guidance. What happens if the audio file is corrupted, the language is unsupported, the model fails to load, or the audio is too long? Errors should explain what went wrong and suggest next steps (e.g., 'Unsupported language. Try list_whisper_models first' or 'File too large (>500MB). Split into chunks.').
No mention of performance characteristics. How long does transcribe_audio take for a 1-hour audio file? Are there timeouts? What is the rate limit? Without this context, LLMs cannot plan efficiently or recognize hung requests.
Model parameter enum has no guidance on which model to choose. E.g., 'tiny' is smallest/fastest, 'large' is most accurate. LLMs need context to pick wisely. 'Supported sizes: tiny (80M, fast), base (140M, balanced), small (240M), medium (769M), large (2.9G, highest accuracy). Default base balances speed and quality.' is actionable.
temperature parameter defaults to 0 but lacks any explanation of its effect. Is it used for sampling? Why does Whisper have temperature if it's typically deterministic? Parameter descriptions should explain when and why to adjust this.