MCP server providing code analysis, security scanning, testing, DevOps, and context memory tools for AI-powered development workflows
This MCP server has 19 tools with widely varying quality. Core issues: (1) Many tools lack complete input schema definitions, parameters are listed but type information is sparse or missing for non-scalar fields like 'metadata' (object type with no nested schema). (2) Descriptions are present but often generic and lack actionable detail for LLM selection, e.g., 'Store a conversation in context memory' does not explain WHEN to use this vs. other context tools. (3) Output schemas are completely undocumented, no tool description specifies what fields are returned, forcing LLMs to guess. (4) Parameter naming inconsistencies and lack of natural-language hints: 'max_turns', 'limit', 'node_types' lack context on valid ranges and formats. (5) No evidence of error handling guidance in any description. (6) Several tools (agentic_team_* and orchestrator_*) appear to wrap multi-step workflows but do not document chaining requirements or output field names needed by downstream tools. Average tool description length ~120 chars (below the 194-char baseline), and NO tools document return structure. This is typical community-grade code, functional but not optimized for agent interaction.
Get team role configuration: roles, agents, and responsibilities.
Execute a task with the agentic team using role-based collaboration. A team of 5 roles (Project Manager, Architect, Developer, QA, DevOps) communicates freely. The Project Manager gates final delivery.
Check agentic team health: status, team validity, agents.
List available agents in the agentic team.
Validate team configuration: check all roles map to available agents.
Generate CI/CD configuration file.
Generate a docker-compose.yml file.
No output schemas documented. 19/19 tools lack documented return types. LLMs cannot infer what fields responses contain, forcing guesswork on downstream tool chaining.
Incomplete input schemas. Many parameters are typed as 'object' (e.g., 'metadata', 'services') with no nested field definitions. This prevents validation and forces LLMs to hallucinate structure.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | F | 46 | 2026-07-28+ | v2 |
Generate a Dockerfile for a project.
Get statistics about the context memory.
Get relevant context for a task.
Log a mistake for learning.
Execute a task through an orchestrator workflow. Routes the task through a sequence of AI agents (e.g. codex -> gemini -> claude) for implementation, review, and refinement.
Check orchestrator health: engine status, agent count, workflows.
List available orchestrator agents with their roles.
List available workflows with their step sequences.
Search context memory using hybrid search.
Store a conversation in context memory.
Store a learned pattern.
Store a task execution in context memory.
Generic, short descriptions lack actionable LLM guidance. Descriptions like 'Store a conversation in context memory' (35 chars) do not explain WHEN to call this tool vs. store_task, what metadata keys are expected, or what the tool returns. Average description length 120 chars vs. baseline 194 chars.
Parameter descriptions are present but lack constraints. 'max_turns' (1-50) and 'limit' have no guidance on typical values or why an LLM would choose one over another. No parameter has enum constraints, ranges, or format hints (e.g., 'ISO 8601 date').
No error handling guidance in any tool description. None of the 19 tools document what happens on failure, how to retry, or what recovery steps are available. LLMs will not know whether to retry or ask the user.
No tool chaining documentation. 'orchestrator_execute' and 'agentic_team_execute' are complex workflows but do not specify what IDs/references must be passed between them, or what the output contains that downstream tools accept. This forces discovery calls.
Ambiguous parameter names lack type suffixes. 'node_types' in search_context is an array but the description does not list valid values. 'services' in generate_docker_compose is an array of strings but does not clarify service name format or available options.
No distinction between read-only and destructive operations in descriptions, even though tool annotations (toolAnnotations=true) are declared. Tools like 'store_conversation', 'store_task', 'log_mistake' modify state but do not declare this in descriptions or use readOnlyHint/destructiveHint in schemas.