An MCP agent system with tool discovery, container-based code execution, and proxy routing to multiple backend MCP servers
This is a meta-tool proxy server that exposes 6 tools for discovering and executing MCP tools. Most tools have descriptions and basic schemas, but critical issues severely limit production readiness: (1) Tool descriptions are present but lack depth about WHEN and WHY to use each tool, missing the LLM-optimized guidance required by pattern:tool-description. (2) Parameter descriptions are thin or missing entirely, e.g., 'code' in container__execute has a description but lacks constraints on what imports are available, timeout behavior, or error modes. (3) Output schemas are NOT documented anywhere, the rubric requires documenting what fields callers should expect, but this server provides no schema for response structures. (4) Security is a critical gap: container__execute and bash__execute are IRREVERSIBLE tools that can modify system state, but lack confirmation workflows, dry-run modes, or permission gates. (5) Error handling is minimal, no guidance on recovery, retryability classification, or actionable next steps. (6) Several tool names violate verb-first conventions: 'search_tools', 'list_tool_names', 'get_tool_definition', 'container__execute', 'bash__execute', and 'container__list_tools' are passable but inconsistent (namespace prefixing on some, not others). The naming inconsistency (search_tools vs get_tool_definition vs container__execute) signals unclear composition. Tool definitions are visible in source (src/mcp-proxy/meta-tool-proxy.ts), so no capping due to inference.
Execute a bash command and return the output. Commands run in the specified working directory.
Execute TypeScript code in an isolated Docker container with access to MCP tools
List all available MCP tools from the container proxy
Get the full definition (name, description, input schema) for a specific tool.
List available tool names. Much faster than search_tools when simple filtering might work. Use get_tool_definition to fetch full tool definitions for specific tools.
Search for available MCP tools using an AI agent to select the most relevant tools. Returns full tool definitions.
No output schemas documented. The rubric requires documenting return types (100% of A+ tools have them). search_tools, list_tool_names, and get_tool_definition must declare what fields they return so agents can plan downstream calls. container__execute and bash__execute must document success/error response structures.
Irreversible tools (container__execute, bash__execute) lack confirmation workflows or dry-run modes. These tools can modify system state but provide no recovery pattern. Per pattern:confirmation-request, destructive operations should support a dry-run or confirmation step to prevent catastrophic errors.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 44 | 2025-06-18+ | v2 |
| 2026-03-09 | D | 52 | - | v1 |
Error handling provides no recovery guidance. Errors must tell the LLM what to do next (pattern:recovery-guide). Currently, if a tool call fails, the agent has no hint about retryability, user-fixable issues, or next steps. All error responses should categorize failure type and suggest recovery actions.
Parameter descriptions lack actionable constraints. 'code' in container__execute is described as TypeScript code but does not specify: What imports/globals are pre-loaded? What happens on timeout? What error format is expected? Per pattern:constrained-input, descriptions must include expected format, range, and allowed values.
Inconsistent tool naming convention. 'search_tools', 'list_tool_names', and 'get_tool_definition' use underscore-separated verb_noun. But 'container__execute', 'bash__execute', and 'container__list_tools' use double-underscore namespace prefixing. Either namespace all tools consistently or use single prefix for discovery tools only. Naming inconsistency confuses LLM tool selection.
No permission gates on sensitive tools. bash__execute and container__execute allow arbitrary code execution. Per pattern:permission-gate, destructive or sensitive tools must verify the calling agent has authority before executing. Implement access control checks.
Tool descriptions lack WHEN/WHY guidance. Per pattern:tool-description, descriptions should explain WHAT the tool does, WHEN to use it, and any prerequisites. 'Search for available MCP tools using an AI agent' does not tell the LLM when search_tools is better than list_tool_names (search_tools uses AI selection, list_tool_names is faster for simple filtering). This distinction is buried in secondary descriptions and forces LLM guessing.