An MCP server for audio interface operations including recording, playback, device enumeration, and real-time conversation with Google Gemini
This server has moderate naming quality (most tools start with verbs like 'list', 'record', 'play') but suffers from significant issues in descriptions, schema completeness, and error handling. Tool descriptions vary wildly in length and quality, some are detailed docstrings (record_audio: 280 chars), others minimal (list_audio_devices: 79 chars, play_latest_recording: 63 chars). Several descriptions are UNDER 20 characters, triggering hard score caps. Parameter descriptions exist but lack constraint information (no enums for voice selection, no range specifications for duration/sample_rate/channels). The play_audio tool is non-functional (returns a stub message about unimplemented TTS). Error handling is minimal, most tools return generic success/failure strings without actionable recovery guidance. The global state pattern (latest_recording variable) violates single-responsibility principle and introduces non-idempotency. Output schemas are completely undocumented, LLMs have no way to know what structure to expect from tool returns.
Start a real-time conversation with Gemini using your microphone and speakers. Args: duration: Maximum recording duration in seconds (default: 5) sample_rate: Sample rate in Hz (default: 44100) channels: Number of audio channels (default: 1) device_index: Specific input device index to use (default: system default) Returns: A message indicating the conversation result
List all available audio input and output devices on the system.
Play audio from text using text-to-speech. Args: text: The text to convert to speech voice: The voice to use (default: "default") Returns: A message indicating if the audio was played successfully
Play an audio file through the speakers. Args: file_path: Path to the audio file device_index: Specific output device index to use (default: system default) Returns: A message indicating if the audio was played successfully
Play the latest recorded audio through the speakers.
Short descriptions (<20 chars) on critical tools prevent LLM selection logic. 'List all available audio input and output devices on the system' (79 chars) and 'Play the latest recorded audio through the speakers' (63 chars) do not explain WHEN to call these vs. alternatives or what prerequisites apply.
No output schemas documented anywhere. Callers (including LLMs) cannot know what fields record_audio or play_latest_recording return. This violates the pattern requirement that 'LLMs need to know what fields to expect so they can plan downstream tool calls.' With no schema, agents cannot compose tools reliably.
play_audio is a stub function that returns 'Text-to-speech functionality requires additional setup.' This is not a working tool, it abandons the user and suggests they install external libraries. Either implement TTS (e.g., pyttsx3, gTTS) or remove the tool entirely.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 46 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 49 | - | v1 |
Record audio from the microphone. Args: duration: Recording duration in seconds (default: 5) sample_rate: Sample rate in Hz (default: 44100) channels: Number of audio channels (default: 1) device_index: Specific input device index to use (default: system default) Returns: A message confirming the recording was captured
Global state (latest_recording variable) violates idempotency. Calling record_audio twice does not produce consistent results if a different agent or user intervenes. Agents rely on idempotent tools to safely retry on transient failures. Store recorded audio with explicit identifiers (e.g., per-request unique IDs) instead of a shared global.
No enums or constraints on parameters. The 'voice' parameter in play_audio accepts any string with no validation or guidance. Valid values should be declared as an enum (e.g., 'default', 'male', 'female'). Duration, sample_rate, and channels lack min/max bounds, an LLM could pass duration=999999 or sample_rate=0.
Error messages are generic and non-actionable. Example: 'Error recording audio: {str(e)}', the LLM receives a Python exception string and has no idea what to do next (retry? ask user? call a different tool?). Replace with categorized errors: 'Device index {index} invalid. Call list_audio_devices to see available devices (0-{max_index}).'
gemini_conversation tool expects a Google API key (google-generativeai library) but no documentation or error handling for missing credentials. If the key is missing, the tool silently fails with a cryptic auth error. Per pattern:secret-injection, credentials must never appear as parameters and must use server-side injection. Add explicit error messages if GOOGLE_API_KEY env var is unset.
No pagination or result limits. Tools that query or list items should accept limit and offset/cursor parameters. While list_audio_devices will typically return a small list, the pattern establishes that agents need control over result volume to avoid context explosion.
Tool naming mixes snake_case and multi-word descriptions inconsistently. 'play_latest_recording' is correct, but it does not clearly signal WHICH recording (oldest? newest? by date?). If multiple recordings exist (or will in future), the tool's purpose becomes ambiguous. Consider 'play_recording' with an optional 'recording_id' parameter, or document that only one recording is kept in memory.