Static source inference · medium confidence · detected: Logging
Deprecated protocol patterns detected
Summary
The mcp-executor defines 4 execution tools with reasonable verb-based naming (execute-python, execute-bash, execute-typescript, execute-go). All tool definitions are directly visible in server_test.go. However, the implementation has critical gaps in schema completeness, parameter descriptions, and output documentation. Schemas show basic parameter structure (code, dependencies, envVars) but lack critical validation rules, constraints, and comprehensive output schemas. Descriptions are present but generic and lack the specificity required for LLM tool selection and error recovery. All tools are WRITE-risk operations but lack explicit recovery guidance or dry-run support. The server provides no documented error handling patterns, no idempotency guarantees, and no per-parameter constraints (e.g., max code length, valid Python versions). Tool names are appropriate verbs but descriptions do not explain when to use each variant or how they differ from alternatives.
Tools (4)
execute-bashwritesource verified52/100
Execute bash scripts with optional package installation support
execute-gowritesource verified52/100
Execute Go code with optional dependency installation support
execute-pythonwritesource verified52/100
Execute Python code with optional dependency installation support
execute-typescriptwritesource verified52/100
Execute TypeScript code with optional npm package installation support
Output schema not documented. Tools execute code and return results, but the response structure is not defined in tool registration. LLMs cannot plan downstream chains or extract specific fields.
No input validation rules or constraints in parameter descriptions. 'code' parameter accepts unbounded strings with no guidance on max length, security restrictions, or sandbox limits. envVars object has no documented value format, character restrictions, or injection prevention.
No error recovery guidance. When code execution fails (syntax error, timeout, missing package), the error response must tell the LLM what to do next. Current descriptions provide no guidance on retryable vs. fatal errors.
Document the complete output schema for each tool. Include: stdout (string), stderr (string), exit_code (integer), duration_ms (integer), success (boolean). Show example JSON response.
Add validation constraints to 'code' parameter: max 50KB, no null bytes, timeouts (30s default, 5m max). Document in description: 'Python code to execute (max 50KB; execution timeout 30s).'
Expand tool descriptions from ~30 chars to 150-200 chars. Include: what the tool does, when to use it instead of alternatives, what output to expect, and any prerequisites. Example for execute-python: 'Execute Python code in an isolated sandbox. Returns stdout, stderr, and exit code. Supports code snippets up to 50KB and stdlib + pip packages (docker mode only). Execution timeout: 30s. Use this to run calculations, transformations, or data processing tasks. For long-running jobs, use execute-bash with async patterns.'
Add explicit mode-based capability statements: 'In subprocess mode: no dependency installation. In docker mode: pip packages (python), apt packages (bash), npm packages (typescript), go get packages (go) supported. Installation persists only within single execution context.'
Document error scenarios in each tool description. Add a brief 'Error Recovery' section: 'If execution fails, check: (1) syntax errors in code, (2) missing dependencies, (3) timeout (re-run with simpler code), (4) sandbox limits (reduce output verbosity).'
Add per-parameter constraints: code (1-50000 chars, no null bytes), dependencies (array of 0-20 items, each 1-256 chars matching package name patterns), envVars (object with 0-50 keys, each key matching [A-Za-z_][A-Za-z0-9_]*)
Spec posture evidence
Inferred effective spec: <=2025-11-25.
Relies on Logging (deprecated) - log to stderr or use OpenTelemetry
Descriptions lack specificity about dependencies behavior. 'Python packages to install via pip' does not explain: does installation persist across calls? Are there size/count limits? What happens if installation fails? Subprocess mode omits dependency support but Docker mode includes it, this distinction is not reflected in the description.
No confirmation or dry-run pattern for destructive/irreversible operations. These tools execute arbitrary code (WRITE risk) which may delete files, modify system state, or send data. No safe-by-default mechanism or user confirmation step is documented.
Parameter descriptions are generic and under-inform LLM usage. 'Python code to execute' does not describe: expected return format, stdout/stderr capture, timeout behavior, or available stdlib. Compare to baseline: avg param description is 72 chars; these are ~20-30 chars.
No documented difference between subprocess and docker execution modes. Tool descriptions do not indicate that dependency installation is only available in docker mode, causing LLMs to attempt unsupported operations or misunderstand capabilities.
envVars parameter lacks type specificity. Described as 'object' but no guidance on key/value format, allowed variable names, or injection prevention. LLMs may pass malformed env structures.
Add idempotency guidance: 'These tools are NOT idempotent. Code that modifies external state (file writes, API calls, DB updates) will have side effects on each call. Agents should verify side effects before retrying.'
For destructive operations, add a recommendation: 'Use dry-run patterns (e.g., Python: preview file writes without executing, or bash: echo commands instead of executing). This prevents accidental data loss when agents retry.'
Document timeout behavior: 'Execution timeout: 30 seconds for subprocess mode, configurable for docker. If code exceeds timeout, execution halts and returns 124 exit code. Long-running tasks should use async patterns or be split into smaller chunks.'
Clarify dependency persistence: 'Installed dependencies persist only within the current execution context. Each invoke() call runs in a fresh environment unless the server is configured with persistent containers.'
Add security guidance in descriptions: 'Code execution is sandboxed. However, avoid executing untrusted code. All parameters are assumed clean; sanitize any user-provided input before passing to this tool.' Reference the pattern:tool-gateway.
Create a usage guide prompt: 'system-code-execution' or similar, documenting best practices for each language (e.g., prefer import statements for dependencies, use print() for output capture, avoid blocking I/O).