MCP server for delegating tasks to specialized AI assistants in Cursor, Claude Code, Codex, Command Code, Google Antigravity, Gemini, GLM, Kimi, Grok, and OpenCode
Single tool server with solid schema definition and comprehensive input validation. Tool naming follows verb-noun pattern (run_agent). Description is well-written and explains the purpose, behavior, and session management. Input schema is fully defined with all parameters typed and described. However, output schema is inferred from interface definitions rather than explicitly documented in the tool registration, and there is no documented error recovery guidance or examples of how session_id reuse works in practice. The tool accepts complex configuration (agent name, prompt, cwd, optional session_id) with good validation logic (max lengths, path traversal checks, format validation), which is above average for community servers.
Delegate complex, multi-step, or specialized tasks to an autonomous agent for independent execution with dedicated context (e.g., refactoring across multiple files, fixing all test failures, systematic codebase analysis, batch operations). Returns session_id in response metadata - reuse it in subsequent calls to maintain conversation context continuity across multiple agent executions.
Output schema not explicitly documented in tool definition. While McpResponseData interface exists in source code, tool consumers cannot see the expected response structure (fields: result, session_id, agent, exit_code, execution_time, status, request_id, session_saved). Tool description mentions 'Returns session_id in response metadata' but the full output structure should be documented in a public schema field.
No error recovery guidance in descriptions. Tool can exit with COMMAND_CODE_MAX_TURNS_EXIT_CODE (8), TIMEOUT_EXIT_CODE (124), SIGKILL_EXIT_CODE (137), SIGTERM_EXIT_CODE (143), or other error states. Description does not explain what these exit codes mean, what caused them, or how to recover. LLM will not know whether to retry, adjust parameters, or request user intervention.
Session management behavior underspecified. Tool description states 'Reuse the returned session_id in subsequent calls to maintain context continuity' but does not explain: How long are sessions stored? What happens if an invalid session_id is provided? Can multiple agents share a session? Does context carry across different agents or only within one agent? This lack of clarity invites misuse.
Inferred effective spec: 2025-06-18+.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 64 | 2025-06-18+ | v2 |
| 2026-03-09 | F | 30 | - | v1 |
No confirmation or dry-run capability for IRREVERSIBLE operations. Tool is marked with Risk=IRREVERSIBLE (agent execution can modify filesystem, run code, etc.). No dry-run mode, no confirmation step before execution. Agents can cause unintended side effects without safety gates.
Agent parameter requires exact name matching. Description states 'Agent name exactly as listed in list_agents resource.' However, no guidance on discovering available agents or handling typos. If user says 'refactor agent' but registered name is 'RefactorAgent', call fails with unclear error.