Multi-agent AI system that orchestrates specialized agents (Architect, Coder, Ops, Reviewer) to build production-ready backend applications. Includes intelligent SuperCoderAgent alternative and integrates MCP tools for code validation, execution, and quality checking.
This server has significant definition quality gaps. While 13 tools are defined with basic input schemas and descriptions, the definitions are inconsistent, lack depth, and violate several critical patterns. Naming is generally acceptable (verb-driven), but descriptions are often too brief (many under 100 chars, well below the 194-char production baseline). Parameter descriptions are sparse or missing entirely for optional parameters. Output schemas are almost entirely undocumented, the tools return Dict[str, Any] with no field specifications. Error handling guidance is minimal. Security-sensitive tools like execute_command and run_python_code lack clear permission declarations and constraint documentation. The server appears to be a multi-agent backend builder, but tool design prioritizes breadth over quality, each tool does too much (e.g., run_python_code runs arbitrary code with only a timeout constraint) and lacks the composition patterns needed for safe agent orchestration.
Main orchestration tool that coordinates all agents to build a production-ready backend from project requirements
Check if all imports in Python code are available or installable
Create a file with validation (syntax checking for Python files, import validation)
Execute terminal commands safely with command sanitization, workspace enforcement, and timeout protection
Generate production-ready code based on the architecture plan, including main application, models, routes, tests, and configuration files
List files in a directory using MCP Filesystem protocol
Design the backend architecture including project structure, dependencies, API design, and database schema
Output schemas are entirely undocumented. Tools return Dict[str, Any] with no field type specifications, making it impossible for LLMs to plan chained calls or extract relevant data. E.g., run_python_code returns {success, stdout, stderr, return_code, execution_time} but LLMs cannot verify these fields exist before using them.
execute_command lacks permission gates and safety constraints. The tool accepts arbitrary shell commands with only a 'test_command' validation function. No documentation of what permissions are required, no rate limiting, no command injection prevention guidance. This is a critical security gap for a production backend system.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 8 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 38 | - | v1 |
Read file contents using MCP Filesystem protocol
Perform comprehensive code review checking code quality, best practices, security, and test coverage
Run comprehensive linting checks using black and flake8 formatters
Execute Python code safely and return execution results including stdout, stderr, and return code
Run pytest on a project and return test results including passed/failed counts
Setup deployment infrastructure including Docker, Docker Compose, CI/CD pipelines, and environment configuration
Validate Python code syntax and detect syntax errors
Write content to a file using MCP Filesystem protocol
run_python_code executes arbitrary Python with minimal constraints. Description lacks warnings about security implications, no mention of sandbox limitations, no guidance on what payloads are safe. For a multi-agent system, this needs explicit 'DESTRUCTIVE' annotation and dry-run capability.
Parameter descriptions are incomplete or missing for optional/advanced parameters. E.g., execute_command's 'timeout' parameter has a generic description ('Command timeout in seconds (optional)') but no guidance on valid ranges, defaults, or what happens on timeout. run_python_code's 'timeout' is identical. These should specify min/max values and timeout behavior.
Tool descriptions are too brief and lack recovery guidance. E.g., 'Execute a terminal command in the workspace' (44 chars) does not explain when to use execute_command vs run_python_code, what 'workspace' means, or what to do if a command fails. Production baseline is 194 chars average; most descriptions here are 40-100 chars.
Error handling lacks actionable recovery guidance. Tools return structured dicts with success/failure but no guidance on what to do next. E.g., if run_tests fails, should the agent fix code and retry, or escalate? If execute_command times out, is it retryable? Error responses must include next-step guidance.
Tool composition is loose. Many tools operate on 'project_path' or 'code', but downstream tools do not clearly document what form code must be in, where project_path points to, or how to chain results. E.g., build_backend and plan_backend_architecture both take requirements but don't explain their relationship or expected sequence.
Parameter types are basic. No enums for constrained inputs (e.g., no limit on command complexity, no allowed-command whitelist). No regex patterns for file paths, no documentation of path constraints, no prevention of path traversal. For a system that manipulates the filesystem and runs commands, this is risky.
No pagination or result limits documented. Tools like get_command_history and get_command_statistics lack limit/offset parameters and max-result guarantees. If an agent calls get_command_history without a limit, could it retrieve thousands of records and blow the context window?
Tool names could be more specific. 'execute_command' is generic, does it run shell, Python, or arbitrary binaries? 'build_backend' is vague, does it scaffold, code-gen, deploy, or all three? Ambiguous names force LLMs to guess intent before reading descriptions.