Engineering-first AI-amplified development methodology with code quality validation, context scoring, multi-agent orchestration, and progressive validation loops
DevMethod is a bash-based MCP server with 21 tools that suffer from critical structural deficiencies. While tool names generally follow verb_noun conventions (validate_project, create_directory_structure), the implementation has severe issues: (1) Tool definitions are inferred from bash scripts rather than explicitly registered via MCP protocol; (2) Descriptions are present but often generic and lack LLM-optimization guidance; (3) Input schemas are visible in the sample Python code but unclear how they map to actual MCP registration; (4) No evidence of error handling that guides LLM recovery; (5) No output schemas documented; (6) Multiple tools combine concerns (e.g., setup_orchestration both initializes AND generates instructions); (7) WRITE operations lack confirmation/dry-run patterns. The codebase shows implementation effort but minimal adherence to agentic tool patterns.
Decompose task into parallel work streams with agent assignments, dependencies, and synchronization checkpoints
Calculate score for a single DevMethod context dimension based on completeness (60%) and quality (40%)
Detect potential file conflicts and overlapping ownership between parallel agents
Verify system has all required tools: git, node.js, npm, python3, pip3, jq, curl, and optional docker
Create DevMethod configuration file at ~/.devmethod/config.yaml with core settings, agent definitions, validation levels, and automation rules
Create DevMethod directory structure at ~/.devmethod with templates, scripts, cache, logs, and projects subdirectories
Define synchronization checkpoints between parallel phases with validation criteria and conflict resolution
Tool definitions inferred from bash/Python scripts rather than explicitly registered via MCP protocol. No visible tool registration mechanism (e.g., Tool objects with names, descriptions, and input schemas). Cannot verify actual MCP schema compliance.
No documented output schemas. LLMs cannot determine what fields to expect from tool responses. Tools like validate_project return dictionaries with 'errors', 'warnings', 'passed' but this is inferred from code, not declared in tool schema.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 38 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Generate parallel work instructions for a specific agent (architect, pm, qa, dev) with file ownership, dependencies, and deliverables
Install Python dependencies (pyyaml, requests, beautifulsoup4, markdown, jinja2) and Node.js global packages
Initialize and monitor real-time progress of parallel agents with status tracking files
Score DevMethod context based on 12 dimensions with completeness and quality metrics
Configure Claude Code integration with DevMethod commands for context gathering and PRP generation
Initialize multi-agent orchestration environment with global state, file ownership matrix, and synchronization structures
Validate a single file length against DevMethod limit of 500 lines
Validate JavaScript/TypeScript files using ESLint with DevMethod configuration for code limits
Level 1 WIRASM validation: syntax, style, and code limits checking with ESLint, TypeScript, Black, Flake8, MyPy, and DevMethod code limits
Level 2 WIRASM validation: unit testing with coverage threshold (80%) and performance budget validation
Level 3 WIRASM validation: integration testing, end-to-end testing, and API contract validation
Level 4 WIRASM validation: production readiness including security audit, deployment validation, and performance testing
Validate entire project against DevMethod standards including file length, function length, cyclomatic complexity, and cognitive complexity limits
Validate Python function lengths and complexity including cyclomatic complexity calculation
Multiple tools combine multiple concerns, violating single-responsibility principle. setup_orchestration initializes state AND generates instructions. create_sync_checkpoints defines checkpoints AND handles conflict resolution. These should be split for agent composability.
WRITE operations (setup_orchestration, create_directory_structure, install_dependencies, create_configuration, setup_claude_integration, generate_agent_instructions, create_sync_checkpoints) lack confirmation, dry-run, or idempotence documentation. High risk of unintended side effects.
No evidence of error handling that guides LLM recovery. Python validator catches exceptions but catches are generic ('Could not analyze'). No categorization of errors as retryable, user-fixable, or fatal. No recovery hints.
Descriptions for setup and orchestration tools are vague (40-50 chars). E.g., 'Initialize multi-agent orchestration environment...' lacks WHEN to use it, WHAT prerequisites exist, and WHAT it returns. Should be 50-200 chars with actionable guidance.
Parameter descriptions are minimal or missing for complex inputs. E.g., calculate_dimension_score accepts 'data' as object but does not specify required fields, structure, or format. score_context accepts 'file' but does not specify JSON structure expected.
Tool names are occasionally ambiguous or generic. 'analyze_task' does not convey decomposition. 'check_requirements' and 'check_conflicts' are generic verbs without clear distinction. Consider 'decompose_task_into_streams' and 'detect_file_ownership_conflicts'.
No idempotence guarantees or retry safety documentation for setup tools. If create_directory_structure or install_dependencies fails mid-way, unclear if safe to retry. No mention of idempotent markers or state tracking.
Tools return raw data structures without field consistency. validate_project returns errors/warnings/passed but unclear if 'errors' is array of objects or strings. No schema ensures downstream tools can rely on predictable field names and types.