Agent Harness for automated generation, deployment, and testing of full-stack codebases. Exposes tools for local specification analysis and planning, document reading, parallel code generation, status polling, output download, and generation retry.
SpecFlow MCP server has 7 tools with mixed quality. Tool names use action verbs (check_, run_, download_, retry_, read_) which is good. However, descriptions vary significantly in completeness and clarity. Most tools have schema definitions with parameters including type and description fields, which is a baseline requirement. Naming is mostly clear except for ambiguous distinction between check_specification_completeness and run_planning (both analysis-type operations). Several critical issues: (1) error handling descriptions are sparse, tools do not explicitly state what errors are retryable, user-fixable, or fatal; (2) schema output documentation is minimal, tools describe what they do but not what fields callers should expect in responses; (3) parameter descriptions lack constraint information (enums, ranges, formats); (4) no evidence of tool annotations (readOnlyHint, destructiveHint, idempotentHint) in the visible schema. The read_document tool is a generic utility that may not fit the specflow-specific domain focus. Tools like run_generation and download_outputs perform state-changing operations but lack explicit confirmation or dry-run patterns. Overall, definitions are functional but fall short of production-grade LLM-optimization standards.
Analyze specification completeness locally. Returns only the agent instruction template (bundled `specflow-analysis` SKILL.md). Your IDE agent reads `spec_dir`, inspects optional brownfield code under `src_dir`, and writes `{outputs_dir}/analysis/specification_completeness.md`. No backend sync, API key, workspace, or generation session. Safe to call repeatedly; does NOT start generation.
Poll the status of an ongoing or completed generation. Returns the current phase, progress, any errors, and agent warnings. Only call when the user explicitly asks 'what's the status?' — do not implement automatic polling loops. Returns immediately with cached status; does not block.
Download and extract generation outputs (generated code, deployment config, test reports) to the local workspace. Only call when the user explicitly requests the artifacts. Automatically finds the generation_id from specflow_session.json if not provided. Downloads a tar.gz archive and extracts it to the project root.
Read and parse a local document file (Markdown, Word, PowerPoint, Excel, PDF, or image). Returns the document content in a structured format, handling text extraction, table parsing, and image-to-text conversion. Supports formats: .md, .docx, .pptx, .xlsx, .pdf, .png, .jpg, .jpeg, .gif, .webp, .bmp, .tiff.
Missing error classification and recovery guidance. Tools like run_generation and retry_generation do not document what errors are retryable vs. fatal, or how LLMs should respond to failures.
Sparse output schema documentation. Tools describe their side effects but provide minimal guidance on response structure, field types, or what downstream tools should expect.
Parameter descriptions lack constraint information. Parameters like spec_dir, outputs_dir, src_dir have descriptions but no explicit format, path validation rules, or error guidance if paths are invalid.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 63 | <=2025-11-25 | v2 |
Resume a failed generation from the last successful checkpoint. Call only if the initial run failed and you want to continue from where coding already started. The user must explicitly request a retry — do not auto-retry on timeout (the operation may have already started). Returns immediately with a new or reused generation_id.
Start full-stack code generation from validated specification and planning files. Uploads spec_dir and planning outputs to the backend, validates them, then runs 2–8 hours of autonomous code generation. Optionally includes deployment and E2E testing. Returns immediately with a generation_id; work runs asynchronously in the background. Emails USER_EMAIL when complete. If local files are missing or invalid, rejects immediately with actionable guidance.
Generate implementation planning template locally. Returns the `specflow-planning` SKILL.md template for the agent to follow. Agent reads the analysis and source code, writes `{outputs_dir}/planning/IMPLEMENTATION_PLAN.md` (and optionally `{outputs_dir}/planning/e2e-test-plan.md` if integration tests are ready). No backend sync, no session. Safe to call repeatedly anytime specs or plans change.
No tool annotations visible (readOnlyHint, destructiveHint, idempotentHint). While tool names hint at safety (check_, read_ are read-only), explicit annotations would improve agent planning.
Confirmation or dry-run pattern missing for destructive operations. run_generation and download_outputs are write operations that could have unintended side effects; no mention of estimation_only or confirmation steps in descriptions.
read_document tool is generic and not domain-specific. It accepts 8 file formats but lacks context for specflow use case. Unclear if this is intended as a discovery/elicitation tool or a utility.