MCP server to enforce development workflow discipline with git integration, testing, documentation, and release management
This MCP server defines 22 workflow tools with substantial gaps in schema completeness, parameter validation, and error handling guidance. While tool names follow action-verb conventions and most have descriptions, critical deficiencies undermine production readiness: (1) Many tools lack explicit input schema validation in the source code provided, schemas are listed in the server submission but not visible in actual handler implementations, making true schema compliance unverifiable. (2) Descriptions are present but often generic (e.g., 'Mark the feature or bug as fixed with a summary of changes' for mark_bug_fixed, 25 chars). (3) Output schemas are not documented in the source, no evidence of structured return type definitions. (4) Error handling lacks recovery guidance; no examples of 'retryable vs. user-fixable' classification. (5) Several tools combine multiple concerns (e.g., run_full_workflow orchestrates 12+ parameters across multiple workflow phases), violating single-responsibility principle. (6) Parameter descriptions lack constraint details (e.g., releaseType accepts 'major|minor|patch' but this isn't stated in the parameter description visible in the source). Tools like perform_release and commit_and_push are high-risk WRITE operations but lack confirmation/dry-run patterns.
Verify that all workflow prerequisites are met before committing
Stage all changes, create a commit, and push to the repository
Mark the current task as complete and reset workflow state for next task
Continue the current workflow from the last checkpoint
Record that documentation has been created or updated
Describe your feature flow using Mermaid diagram syntax
Mark that tests have been created for the changes
run_full_workflow tool combines 12+ workflow phases (testing, documentation, commit, release) into a single orchestration call with 14 parameters. This violates single-responsibility principle and reduces agent composability. An LLM cannot conditionally skip phases or handle mid-workflow errors granularly.
Output schemas are not documented in the source code. Tools return data structures, but no schema documentation is visible, no @returns blocks, no TypeScript interfaces exported, no example responses. LLMs cannot predict return fields for downstream tool chaining.
High-risk WRITE operations (commit_and_push, perform_release) lack dry-run or confirmation patterns. An agent can irreversibly push code or trigger a major release without a safety gate. No evidence of _confirmBefore or equivalent guardrail.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 47 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Abandon the current task and reset workflow state
Force complete a task even if workflow steps are incomplete
Get the current workflow status and phase
Mark the feature or bug as fixed with a summary of changes
Execute release command, create version tags, and push tags to repository
Generate a summary of the project workflow and status
Get structured project summary data as JSON
Get database-backed project summary and workflow history
Restart the workflow from the beginning for the current task
Execute the complete workflow in a single orchestrated call
Record test results and run test command
Skip the release phase with a documented reason
Skip the testing phase with a documented reason (for manual QA)
Declare what you're coding by starting a new task with a description and type
View completed workflow history and summary data
Parameter descriptions lack constraint details. E.g., releaseType is described only as 'Release type', the enum values (major, minor, patch) are not stated in the description, forcing LLMs to guess or require clarification.
Descriptions are generic and under 50 characters for many tools (create_tests: empty, mark_bug_fixed: ~25 chars, project_summary: ~45 chars). Production-grade descriptions should be 50-200 chars and include WHEN to call, WHAT it does, and any prerequisites.
Error handling is not visible in the source code. No evidence of retryability classification, recovery guidance, or actionable error messages. When a tool fails (e.g., git push rejects due to conflicts), there is no guidance for the LLM on what to do next.
Three tools query project summary with overlapping functionality (project_summary, project_summary_data, project_summary_db). This creates LLM decision paralysis and reduces composability. Should consolidate to a single canonical tool with optional output format parameters.
No pagination support visible in list-like tools (view_history accepts 'limit' but no offset/cursor, no documentation of total count or next_cursor return field). Large result sets could exhaust context window.
Tool names like skip_tests, skip_release use imperative verbs but are read-only record-keeping operations. Names should reflect this (e.g., mark_tests_skipped, record_release_skip) to avoid ambiguity about destructiveness.
Stateful workflow assumptions: Tools like continue_workflow and rerun_workflow rely on prior state (current task, checkpoint). No documentation of what state is required or how to initialize it. If called out-of-order, they will fail opaquely.