Give your AI agents the ability to listen. Microphone capture and speech-to-text tools for MCP-compatible agents.
mcp-listen demonstrates strong tool naming, comprehensive parameter descriptions, and explicit input schemas. All three tools are action-verb named and appropriately scoped. Schemas are complete with types and ranges. However, output schemas are not documented in the code, and error handling guidance is not visible in the definitions. Tool descriptions are detailed (100-250 chars), exceeding baselines but appropriately pitched for voice/audio domain complexity. Parameter descriptions are uniformly present and detailed, explaining ranges, defaults, and dependencies clearly.
Record audio from the microphone for a specified duration, or until the speaker stops talking with stop_on_silence, and save as a WAV file. Returns the file path and metadata.
List available audio input devices (microphones) on this machine. Each device has a numeric index, a human-readable name, and a stable id. Prefer the id when selecting a device: indexes can shift when devices are added or removed, and names are not unique.
Record audio from the microphone, transcribe speech to text using local whisper.cpp, send the transcription to a local Ollama LLM, and return the response. Recording stops automatically when the speaker stops talking. Fully offline.
Output schemas not documented in tool definitions. LLMs cannot infer response structure, field names, or data types returned by capture_audio or voice_query.
No error handling or recovery guidance in tool descriptions. Tool descriptions do not explain what failures look like (e.g., 'no audio devices found', 'Whisper model not available', 'Ollama connection failed') or how LLMs should respond.
voice_query parameter relationships not fully documented. The dependency between stop_on_silence and silence_ms is stated, but the interaction between duration_ms, stop_on_silence, and the 10-second fallback condition is complex and could confuse LLMs about expected behavior in edge cases.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | B | 76 | 2026-07-28+ | v2 |
capture_audio and voice_query accept overlapping device parameters (duration_ms, device, stop_on_silence, silence_ms). While not strictly a duplication issue, the descriptions should clarify when to use capture_audio (audio-only) vs voice_query (with transcription/LLM).