MCP server that hosts isolated Strands AI agents in Docker containers, managing agent lifecycle, message dispatch, and agent communication through FastMCP
Mixed quality across two distinct tool groups. Agent management tools (send_message, get_messages, list_agents, stop_agent) have good descriptions and parameters with clear purpose and AWS configuration options. GitHub tools (create_issue through reply_to_review_comment) have reasonable descriptions but lack consistent schema visibility. Parameter descriptions are present but inconsistent in depth. No tool annotations (readOnlyHint, destructiveHint, idempotentHint) are evident. Error handling guidance is minimal. Output schemas are not documented in the definitions provided. The server splits responsibilities appropriately between agent orchestration and GitHub operations, but individual tools lack LLM-optimized descriptions and comprehensive error recovery guidance.
Adds a comment to an issue or pull request.
Creates a new issue in the specified GitHub repository.
Creates a new pull request.
Gets details of a specific GitHub issue.
Gets all comments for a specific GitHub issue.
Get the latest messages from an agent's conversation history. Returns large message payloads that consume context window. Should only be called once when user explicitly asks to check on an agent's progress, not immediately after send_message or in polling loops.
No tool annotations present. Tools lack readOnlyHint, destructiveHint, and idempotentHint metadata. LLMs cannot determine which tools are safe to retry, which modify state irreversibly, and which are read-only without analyzing descriptions.
Output schemas not documented. Tool descriptions do not specify what fields are returned, their types, or structure. LLMs cannot plan downstream tool chains or extract relevant data without documented output schemas.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 59 | <=2025-11-25 | v2 |
| 2026-03-09 | D | 56 | - | v1 |
Gets review threads and comments for a specific pull request using GitHub GraphQL API.
Gets details of a specific pull request.
List all agents and their current status.
Lists issues from the specified GitHub repository.
Lists pull requests from the specified GitHub repository.
Reply to a specific review comment on a pull request.
Send a message to an agent (fire-and-forget). Creates the agent if it doesn't exist. This returns immediately after dispatching the message. The agent processes the message in the background. Use get_messages to check for the response.
Stop an agent's Docker container. The agent's data persists and can be restarted via get_messages with auto_restart=True or by sending a new message.
Updates an issue's title, body, or state.
Updates a pull request's title, body, or base branch.
Error handling lacks recovery guidance. Tool descriptions do not tell LLMs what to do when calls fail: should they retry? Ask the user? Call a different tool? Raw errors provide no actionable path forward.
GitHub tools lack fallback and error context. When 'repo' is missing and GITHUB_REPOSITORY env var is not set, error messages do not suggest how to provide the repo. When a repo is invalid, errors do not suggest valid alternatives.
GitHub token exposure risk. Tools accept optional 'repo' parameter and fallback to GITHUB_REPOSITORY, but token retrieval from CONTAINERIZED_AGENTS_GITHUB_TOKEN env var is implicit. Token should never leak into tool parameters or logs. Verify token is server-side injected only.
GitHub tool descriptions are generic and under-optimized for LLM selection. 'Updates an issue's title, body, or state' does not explain WHEN to use this vs other tools, WHAT format is expected for 'state' (enum: open|closed), or WHY it's different from create_issue. LLM-optimized descriptions should be 50-200 characters and include intent and constraints.
Parameter constraints not enforced in descriptions. 'state' parameter in update_issue and update_pull_request should document valid values (open|closed). 'count' in get_messages should document min/max (e.g. 1-100). Without constraints, LLMs guess wrong and fail.
Pagination strategy undefined for list tools. list_agents, list_issues, list_pull_requests, get_issue_comments do not document whether they support pagination, what limit applies, or how to fetch additional results. Large result sets risk context window exhaustion without documented pagination.