MCP server for speech-to-text and text-to-speech functionality with voice conversation support, project management, and Claude integration
Server has 15 tools with reasonable naming conventions and mostly present descriptions, but exhibits significant schema quality issues and inconsistent parameter documentation. All tools start with action verbs (speak, listen, todo_*, project_*, process_*, claude_chat), which is good. However, parameter descriptions are often minimal or absent, input schemas lack proper type constraints for most numeric/enum fields, and output schemas are completely undocumented. Error handling guidance is minimal. The server demonstrates basic tool composition (todo_* family, project_* family, process_* family) but lacks idempotency annotations and input validation rules. Descriptions are reasonably present (avg ~120 chars, within the 10-1024 baseline) but could be more prescriptive about when to use each tool and what to do on failure.
Run Claude CLI and capture its output. Automatically uses linked project context if available. This allows the server to orchestrate Claude Code instead of requiring the user to run it separately.
Capture and transcribe voice using Google Gemini STT with VAD and chunking. Uses Opus 16kHz mono for optimal bandwidth and speed.
Get the output (stdout/stderr) from a running or completed process.
Start a project process (dev server, build, test). Automatically uses configured project commands or allows custom shell execution.
Get the status of a running process.
Stop a running project process by ID.
Output schemas completely undocumented. No tool declares what it returns, LLMs cannot plan downstream calls or extract data fields. This forces agents to reverse-engineer response structure via trial and error.
Numeric and enum parameter constraints lack validation bounds and format descriptions. E.g. 'vadThreshold' (0-1) and 'minSilenceMs', 'maxUtteranceMs' have no documented min/max values. LLMs may pass unbounded or invalid values.
Minimal parameter descriptions. E.g. 'todo_delete' accepts only 'id' with description 'TODO ID to delete', no guidance on what happens on retry, no error recovery hints. 'project_get' returns an object but no schema/description of what fields exist.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 53 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 47 | - | v1 |
Get the currently linked project for this conversation.
Link a project to this conversation. Initializes project context with path, commands, and settings. Use this at the start of a project to configure dev server, build, and test commands.
Unlink the project from this conversation.
Convert text to speech using ElevenLabs with optional SSML enrichment via OpenAI. Audio quality optimized for speed (16kHz, lower latency).
Archive a completed TODO item. Use when cleaning up completed tasks.
Create a new TODO item. Use when the user requests a task to be tracked, or when you identify action items during conversation.
Delete a TODO item.
List TODO items with optional filters. Use when you need to see task status or get project context.
Update an existing TODO item. Use when task progress changes, or when marking tasks as in-progress, blocked, or completed.
No error handling guidance. Tools provide no recovery hints when failures occur. E.g. 'todo_create' fails silently; no guidance on retry vs user intervention. 'process_start' may fail if command is invalid, no description of what to do.
Destructive/write operations lack confirmation or dry-run support. 'todo_delete' and 'process_start' (especially 'custom' mode) can permanently alter state without safeguards. No idempotent annotations.
Missing result field chaining references. E.g. 'todo_create' likely returns a todo_id, but no response schema shown. If 'todo_update' is called next with that ID, the schema must guarantee the field name matches ('id' vs 'todo_id' vs 'taskId'). Broken chains force extra lookups.
Ambiguous parameter dependencies not documented. E.g. 'process_start' has 'processType' enum with 'custom', when selected, 'command' becomes required. This conditional requirement is not stated in either parameter description.
No pagination documented for 'todo_list'. If many TODOs exist, the tool may return thousands of items, bloating context. No 'limit', 'offset', or 'next_cursor' parameters visible.
Generic tool descriptions lack prescriptive 'when to use' guidance. E.g. 'speak' says 'Convert text to speech' but does not explain when to call it vs fallback behavior, or what happens if audio generation fails.