Spectr MCP server — turn a screen recording into a production-ready spec.md, from inside Claude Code / Cursor / any MCP client.
Spectr MCP exposes a single tool 'generate_spec' with a detailed docstring and well-structured input schema. The tool name is clear and action-oriented (verb_noun), and the description is comprehensive (350+ chars covering WHAT it does, WHEN to use it, and TYPICAL runtime). All input parameters have type declarations and descriptions. However, the output schema is not formally documented, the description says the tool returns an ~80-150KB markdown document inline OR a short summary if output_path is provided, but there's no structured output definition (e.g., no explicit fields like 'status', 'spec_content', 'file_path'). This prevents LLMs from reliably parsing the response and chaining to downstream tools. Additionally, error handling is not explicitly documented, the tool description does not clarify what happens on invalid inputs (e.g., missing ffmpeg, invalid MP4 path, Anthropic API failures, network timeouts). The description also does not address which errors are retryable vs. fatal, and no recovery guidance is provided. For a tool that runs 5 - 10 minutes and can fail for many reasons (bad input, missing dependencies, API quota, network), the lack of error classification is a significant gap. The schema itself is well-formed with appropriate types (string, dict, int) and sensible defaults (max_frames=20, your_app_name defaults to reference_app).
Generate a production-ready spec.md from a screen recording. The spec is a structured 7-section markdown document (~80-150KB) covering app overview, navigation architecture, screen specifications, component library, design system (with exact hex/px/weight values), implementation notes, and a Claude Code prompt the developer can paste to build the clone. Targets Expo SDK 54 / React Native / iPhone 15 baseline. Pipeline runs on the user's Claude subscription via the `claude` CLI by default (no API key needed). If ANTHROPIC_API_KEY is set in the env, uses the SDK path instead. Typical run: 5–10 min per MP4.
Output schema not formally documented. Description says 'returns full ~80-150KB spec content inline' OR 'short summary' if output_path is set, but no structured output definition (fields, types, response structure). LLMs cannot reliably parse the response or chain to downstream tools.
Error handling and recovery guidance missing. Tool can fail in many ways (invalid MP4, missing ffmpeg, Anthropic API errors, timeouts during 5 - 10 min pipeline run), but description does not specify which errors are retryable, what the error response looks like, or what the LLM should do next. No error classification (retryable vs. user-fixable vs. fatal).
Long-running operation (5 - 10 min) lacks explicit progress/heartbeat documentation in tool description. Code shows heartbeat logic (HEARTBEAT_INTERVAL_S = 25) in server.py, but tool docstring does not explain to LLM that it will receive progress notifications during execution. Unclear whether LLM should expect multiple intermediate responses or a single final one.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | D | 59 | <=2025-11-25 | v2 |
Parameter 'brand_colors' is typed as 'dict' with no value constraints. Description gives example '{"primary": "#FF5722"}' but does not enforce hex color format, required keys, or limits on dict size. LLMs may pass invalid colors or arbitrary key-value pairs.