Static source inference · medium confidence · evidence: Streamable HTTP
Current-spec patterns detected
Summary
The server defines 2 tools but has severe quality gaps. echo_tool is trivial (READ_ONLY, minimal utility). instruct-developer is the substantive tool but has a dangerously vague description, no input schema validation, missing parameter type declarations, and critical security issues (arbitrary shell command execution, plaintext secret handling, destructive file operations without safeguards). The description is too long (177 chars) and lacks clarity on what happens when called. No output schema is documented. Error handling exists but is weak, errors are human-readable strings, not structured recoverable responses. No tool annotations (readOnlyHint/destructiveHint) despite one tool being READ_ONLY and the other being highly DESTRUCTIVE. Parameter descriptions are minimal or absent. No pagination, batching, or composition support.
When you need to shallow clone a Git repository, optionally limiting to a specific folder and its descendants, then read its system prompt and agent config and run Codex CLI accordingly.
Tool name 'instruct-developer' violates verb_noun convention and is ambiguous. Does not start with a clear action verb (should be 'clone_and_instruct', 'run_codex', or split into separate tools). Name does not convey the destructive nature (Git clone, file read, Docker execution, Codex invocation).
instruct-developer description is vague and does not state what happens: 'When you need to shallow clone...'. Does NOT explicitly say 'This tool clones a Git repository, reads configuration files, executes Codex CLI via Docker, and runs arbitrary AI-generated code.' Agents cannot infer the destructive consequences. Missing WHEN to use it, prerequisites (Git, Docker, OpenAI API key, .agent/system.md structure must exist), and what is returned.
No input schema visible for instruct-developer. Code shows @mcp.tool('instruct-developer') decorator with function parameters (repository: str, request: str, folder: str = '/') but no explicit schema registration, constraints, or validation. Parameter types are inferred from Python signatures, not declared as JSON Schema. No enum for 'folder', no regex for 'repository' URL validation, no length limits. This violates the pattern:constrained-input rule.
Recommendations
Rename 'instruct-developer' to a verb-first name such as 'run_codex_instruction', 'clone_and_execute_codex', or split into: 'clone_repo_sparse', 'read_agent_config', 'run_codex_instruction' (one tool per action).
Rewrite the tool description to be explicit and LLM-actionable: 'Clones a Git repository (with optional sparse-checkout), reads configuration from .agent/system.md and .agent/agent.json, and executes a Codex CLI instruction via Docker. This tool modifies local files in a temp directory and invokes an external AI model (Codex) which may generate and apply code changes. Requires: Git, Docker daemon, OPENAI_API_KEY environment variable. Use when you need to run AI-driven code generation on a specific repository. Returns the raw Codex CLI stdout/stderr.'
Add JSON Schema input validation. Declare repository as a string with a URL format and optional regex pattern (e.g., '^https?://(github\.com|gitlab\.com)/'). Declare request as a string with minLength=1 and maxLength=2048 to prevent hallucinated instructions. Declare folder as a string matching '^/?[a-zA-Z0-9/_.-]*$' to block path traversal attempts.
Add detailed parameter descriptions: 'repository (required, string, URL format): Git repository URL to clone. Must be a publicly accessible HTTPS URL. Examples: https://github.com/owner/repo, https://gitlab.com/group/project. Do not pass local file paths or SSH URLs.' 'request (required, string, 1-2048 chars): Instruction to send to Codex CLI for code generation. Be specific (e.g., "Write a Python function to sort an array") rather than vague. Codex will attempt to implement the request and may modify files in the cloned directory.' 'folder (optional, string, default '/'): Relative path within the cloned repository to focus on (e.g., 'src/main' limits sparse-checkout to src/main/ and descendants). Must be a valid relative path with no ../ sequences.'
Parameter descriptions are missing or inadequate. 'folder' is described as 'Optional specific folder path within repository (default: '/')' but does not explain the security risk (path traversal), the format (relative vs absolute), or that it must exist in the cloned repo. 'request' is described as 'Request/instruction to send to Codex CLI' but does not explain what Codex CLI is, what instructions it accepts, or that arbitrary prompts can lead to arbitrary code generation and execution.
Secrets exposed in environment variable handling. Code reads OPENAI_API_KEY from environment and also from .env file (line: 'OPENAI_API_KEY='), then passes it as -e OPENAI_API_KEY={openai_api_key} to Docker subprocess. If the tool result or any debug output is logged, the API key leaks into agent traces. The .env file read is particularly dangerous, credentials in files accessible to the process are a compliance violation.
No permission checks or rate limits. Tool can be called with any repository URL, any request string. No validation that the agent/user has authority to clone, read .agent/, or execute Codex. No guards against infinite loops (agent asks Codex to write a tool that calls instruct-developer, which calls Codex again). Rate limiting is absent, a runaway agent can spawn unlimited Docker containers.
Command injection vulnerability. Code constructs a subprocess call: docker_cmd = ['docker', 'run', '--rm', '--tty', '-v', f'{work_dir}:/workspace', '-e', f'OPENAI_API_KEY={openai_api_key}', ...]. If repository URL or request string contain shell metacharacters, and if subprocess.run were changed to shell=True (or if the string were concatenated), injection is trivial. Even with shell=False, the folder parameter (read from user input) is used in git sparse-checkout without validation, a malicious path like '../../../etc/passwd' could traverse.
No output schema documented. Function returns a string (either result output, error message, or formatted info). Agent has no way to know the structure of success vs failure, or what fields to expect. If Codex CLI stdout is returned verbatim, it may be gigabytes of logs, exceeding context limits. No pagination, no structured fields.
Error responses are weak and non-actionable. Examples: 'Failed to clone {repository}: {e}' (raw exception, agent cannot act), 'Failed reading system prompt: {e}. Latest commit: {latest_commit}. items ({len(items)}) are: ...' (human-readable but unstructured, agent cannot parse). No classification (retryable vs fatal), no recovery guidance. An LLM seeing 'Failed reading system prompt' cannot decide whether to try another repo or ask the user.
No confirmation or dry-run mode for a destructive operation. Tool clones repos (creates temp files), reads files, executes Codex (which generates and may apply code changes), and returns results. An agent mistakenly passing a wrong request or repository URL causes unrecoverable side effects (Codex may have already modified the cloned repo). No way to preview or rollback.
Tool combines too many concerns: Git clone (with sparse-checkout logic), file I/O (read .agent/system.md and agent.json), JSON parsing, environment variable resolution, Docker subprocess spawning, and Codex CLI invocation. This violates the pattern:tool rule (one tool = one action). Should be split into separate tools: clone_repo_sparse, read_agent_config, run_codex_instruction.
Missing tool annotations. The FEATURES section marks toolAnnotations=false, confirming no readOnlyHint/destructiveHint/idempotentHint. echo_tool should have readOnlyHint=true. instruct-developer should have destructiveHint=true. This prevents the MCP client from applying appropriate caution (e.g., requiring user confirmation before executing destructive tools).
No audit trail or logging. Code prints to stdout (print statements) but does not log to a structured format (JSON, syslog) with timestamps, user IDs, parameters, and outcomes. Agent-initiated Git clones and code execution must be traceable for compliance and incident response.
Timeout exists (600s = 10 min) but is not documented in the tool description. Agent has no way to know the operation may hang for 10 minutes. If a Codex invocation exceeds the timeout, the error message is likely a generic subprocess.TimeoutExpired exception, which does not guide the agent to retry with a different strategy.
No validation of repository URL format. Code passes repository string directly to git clone without checking if it is a valid URL, a local path, or a potential injection vector. No allowlist of trusted repositories.
Temporary directory cleanup is incomplete. Code creates temp_dir = tempfile.mkdtemp() and removes it with subprocess.check_call(['rm', '-rf', temp_dir]) inside a try block. If the tool succeeds, cleanup happens. But if an exception is raised after temp creation, the directory is not guaranteed to be removed (finally block missing). Leaks disk space and potentially sensitive cloned code.
instruct-developer
Inject OPENAI_API_KEY via secure secret injection, not environment file read. Remove the .env file parsing code entirely. Use a vault-based approach (e.g., AWS Secrets Manager, HashiCorp Vault) or require the calling agent/system to provide the key via a secure server configuration, never as a tool parameter.
Add permission gating. Before executing, verify the calling agent/user has authority to: (1) clone from the given repository URL (allowlist check), (2) read .agent/ config files (permission scope validation), (3) execute Docker and Codex (rate limit check). Return a clear error if unauthorized: 'User does not have permission to clone from {repository}. Authorized domains: github.com, gitlab.com.'
Add input validation and sanitization. Validate repository URL format (HTTPS only, matches allowlist). Validate folder path (no ../ , no absolute paths, no symlinks outside repo). Validate request string (no null bytes, reasonable length). Use shlex.quote() if converting to shell strings. Use pathlib.Path.resolve() to canonicalize folder paths and detect traversal attempts.
Document the output schema in the description. Example: 'Returns a plain-text string containing the Codex CLI stdout and stderr combined. On error, returns a string starting with "Error:" followed by a brief description and recovery hint (e.g., "Error: system.md not found. Ensure .agent/system.md exists in the repository root.").' Then implement structured output: return a JSON object {"success": bool, "output": str, "error": str, "executed_at": ISO8601_timestamp, "duration_ms": int}.
Implement structured error responses with recovery guidance. Examples: On git clone failure: {"success": false, "error": "clone_failed", "message": "Failed to clone from {url}: {reason}. Verify the URL is accessible and public, then retry.", "retry_safe": true}. On missing .agent/system.md: {"success": false, "error": "config_not_found", "message": "system.md not found in .agent/ directory. Available files: {listed files}. Ensure the repository contains .agent/system.md before retrying.", "retry_safe": false}.
Add a dry-run mode. New optional parameter: 'dry_run' (boolean, default false). When true, clone and read config, but do not invoke Codex. Return the config and request that would have been executed, so the agent can preview before committing. If dry_run succeeds and agent confirms, agent calls tool again with dry_run=false.
Add tool annotations. For echo_tool, register with readOnlyHint=true (safe to call multiple times). For instruct-developer (or its split variants), register run_codex_instruction with destructiveHint=true (may modify files and invoke external services). This signals to the MCP client that these are risky operations.
Add structured audit logging. Log every invocation to a file or centralized logging service (e.g., syslog, CloudWatch) with: timestamp (ISO8601), caller_id (agent identifier or user), repository_url, request, folder, outcome (success/failure), duration_ms, error_code (if any). Do not log secrets. Example: {"timestamp": "2025-01-15T14:32:10Z", "caller": "agent-abc123", "action": "run_codex_instruction", "repository": "https://github.com/example/repo", "success": false, "error_code": "config_not_found", "duration_ms": 1250}
Document the timeout (600s) in the tool description and add a retry-with-backoff pattern to the error message. If subprocess.TimeoutExpired occurs, return: {"success": false, "error": "timeout", "message": "Codex instruction exceeded 600-second timeout. Try with a simpler request or a more targeted folder path. Retry with exponential backoff.", "retry_safe": true, "suggested_backoff_seconds": 30}
Refactor temporary directory cleanup to use a context manager or finally block to guarantee deletion even on exception. Example: try: temp_dir = tempfile.mkdtemp(); ... finally: shutil.rmtree(temp_dir, ignore_errors=True). This prevents disk leaks.
Split the monolithic instruct-developer into three focused tools: (1) 'clone_repository_sparse(repository, folder)', clones and returns the cloned_path and HEAD commit hash. (2) 'read_agent_config(repo_path)', reads .agent/system.md and agent.json, returns structured config. (3) 'run_codex_instruction(repo_path, instruction, model_id)', executes Codex and returns structured output with success/output/error/duration. Each tool can be tested, versioned, and composed independently.
Add rate limiting and quotas. Track invocations per agent/user over a time window (e.g., max 10 calls/hour, max 5 concurrent Docker containers). Return a 429-like error if exceeded: {"success": false, "error": "rate_limit_exceeded", "message": "Agent has reached the limit of 10 run_codex_instruction calls per hour. Retry after 300 seconds.", "retry_after_seconds": 300}
Validate and allowlist repository URLs. Maintain a list of approved Git providers (github.com, gitlab.com, etc.) or approved organization URLs. Reject unauthorized URLs immediately with a clear error. This prevents agents from cloning arbitrary, potentially malicious, repositories.
Add logging of secrets in a way that preserves compliance. If OPENAI_API_KEY must be logged (for debugging), hash it or log only the last 4 characters (e.g., 'sk-***...abcd'). Never log the full key. Rotate keys frequently and use short-lived tokens when possible.