A toolkit for running Robot Framework data-driven tests with AI Agent capabilities using Codename Goose and Ollama, supporting both local and Docker-based execution with MCP extensions
This MCP server has significant definition quality issues across all three tools. Tool names are action-oriented (send_*), but descriptions are brief and lack critical context about prerequisites, return values, and error conditions. Input schemas are present with type information, but parameter descriptions are minimal and don't explain constraints, expected formats, or dependencies. No output schemas are documented. The tools expose shell command execution to external systems (Docker containers, local Ollama), which carries substantial security and reliability risks that are not addressed in the definitions. Parameter names like 'first_extension' and 'second_extension' are unclear, the descriptions don't explain what these extensions are, where they come from, or what happens if they're invalid. The code shows subprocess calls with shell=True (tool 2) which is a security anti-pattern, but this risk is not surfaced in the tool descriptions to guide LLM usage safely.
Use the 'goose run' command to send a prompt message to a specific Goose AI Agent running in a specific Docker container.
Use the 'goose run' command to send complex instructions through a markdown file to a local Goose AI Agent running on a local machine that is also running Ollama.
Use the 'goose run' command to send a prompt message to a local Goose AI Agent running on a local machine that is also running Ollama.
Missing output schema documentation. No tool documents what fields are returned, making it impossible for LLMs to plan downstream operations or extract structured data from results.
Parameter descriptions are minimal (under 40 chars). 'first_extension' and 'second_extension' lack context about what these extensions are, where they originate, valid values, and consequences of invalid input. The description 'First MCP extension to load with goose run' does not explain what an MCP extension is, how to name it, or what happens if it doesn't exist.
No error handling guidance. Tool descriptions do not explain what happens on failure (e.g., container not found, goose command fails, instruction file missing), what errors are retryable, or what the LLM should do next. Code shows subprocess.run(..., check=True) which will raise on non-zero exit, but no exception handling is documented.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 50 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 38 | - | v1 |
Incomplete tool descriptions (91 chars, 113 chars, 156 chars). These descriptions lack context on when to use each tool, prerequisites (e.g., is Docker running? Is Ollama installed?), and what the output represents. They assume deep knowledge of Goose AI Agent architecture.
No parameter constraints (enums, min/max, regex patterns). The 'container_name' parameter accepts any string but should probably be an enum of running containers. The 'prompt_message' and 'instruction_file' have no length limits, format specifications, or validation guidance. Free-form strings invite invalid input.
Security risk: shell=True in tool 2. The code `subprocess.Popen(ollama_runner, shell=True, ...)` is a command injection vulnerability. An LLM passing a prompt_message containing shell metacharacters (e.g., `; rm -rf /`) could execute arbitrary commands. The tool description does not warn about this or document input sanitization.
Ambiguous parameter semantics. For send_docker_ai_agent_prompt_message, what are 'first_extension' and 'second_extension'? Are they file paths, package names, URLs? The hardcoded '--with-builtin' in the code is not reflected in the schema, making the tool's actual behavior opaque to the LLM.
No idempotency or side-effect clarity. The descriptions do not state that these tools are stateful (they execute code, modify external systems). Per pattern:command-tool, tools that modify state must be explicitly marked as such so LLMs know they cannot be freely retried.