Local, offline audio transcription MCP server and CLI (whisper.cpp) with WhatsApp chat-export support
Harken is a well-structured audio transcription MCP server with three focused tools. Strengths: all tools have detailed descriptions (200+ chars each), complete JSON Schema input definitions with proper type constraints and regex patterns, explicit tool annotations (readOnlyHint, openWorldHint), and documented output schemas. The tool naming is clear and verb-based. Weaknesses: parameter descriptions are somewhat brief (not all include format/constraint details explicitly in the description text, despite being in the schema), and the tool set is narrowly domain-specific without broad composability patterns. All three tools are read-only operations with clear single responsibilities. Evidence: transcribe_file uses {type:string, description:'Absolute path...'} with schema visibility; transcribe_whatsapp_export includes regex pattern validation for date fields; output schemas for all tools are formally declared with detailed property descriptions.
Transcribe a local audio file (opus, ogg, mp3, m4a, wav, flac, mp4, webm, ...) fully offline via whisper.cpp. Model and language are fixed by the server's --model/--lang flags. Returns the transcript text.
Report this server's fixed configuration and model cache state without transcribing anything: model name, whether the model file is already cached locally (path and size in bytes) or the first transcription call would have to download it first (~466 MB for the default 'small'), the startup warm-up's state ('downloading' means a call waits on a download already in flight; 'idle' means none is running), and whether the warm-up failed.
Extract and transcribe every voice-note attachment in a WhatsApp chat-export .zip (iOS or Android), optionally restricted to an inclusive date range. Fully offline. Returns one line per voice note, prefixed with date, time and sender.
Parameter descriptions lack explicit constraint details in prose. Schema includes regex patterns and types, but descriptions like 'Absolute path to the audio file' do not restate the constraint (e.g., 'must be an absolute path, 1 - 4096 chars'). LLMs cannot read JSON Schema, they rely on description text to understand constraints.
transcribe_status parameter documentation is empty (no parameters required, but description says nothing about what happens when called with an empty object). While the tool description is clear, adding a note like 'Takes no parameters; returns the server's configuration and model cache state' would improve clarity.
Output schema for transcribe_whatsapp_export documents that messages[] items are heterogeneous (some with text/duration, some with error), but the schema's 'required' array only lists four fields without clarifying which combinations are valid. An LLM may not anticipate missing text fields in error cases.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 65 | 2025-06-18+ | v2 |