One brain, many arms — spawn multiple specialized Claude Code agents as MCP servers
claude-octopus exhibits significant gaps in definition quality. Tool descriptions are present but often generic or incomplete. Critical issues include: (1) missing input parameter descriptions across multiple tools (e.g., claude_code_timeline has offset/limit/run_id without explaining what they control); (2) incomplete parameter type information in visible schemas; (3) lack of output schema documentation for any tool; (4) no enum constraints for known value sets (e.g., status parameter in claude_code_sessions, permission_mode in create_claude_code_mcp); (5) naming follows verb_noun pattern acceptably but descriptions don't explain error recovery, prerequisites, or downstream dependencies. Tool descriptions average ~150 chars (within the 34-392 baseline range) but lack LLM-specific guidance on WHEN to call each tool and WHAT to do if it fails. The factory tool (create_claude_code_mcp) exposes permission_mode and model as free-form strings instead of enums. No tool documents what fields to expect in responses, and no output pagination patterns are declared despite tools like claude_code_timeline accepting pagination parameters.
Send a task to an autonomous Claude Code agent. It reads/writes files, runs shell commands, searches codebases, and handles complex software engineering tasks end-to-end. Returns the result text plus a session_id for follow-ups via claude_code_reply.
Continue a previous Claude Code session using its session_id. Allows follow-up prompts to the same agent context.
Generate an HTML report of agent execution history. Includes metrics, transcripts, and performance analysis.
List all saved sessions with their metadata and status. Allows resuming past agent invocations.
Query the cross-agent execution timeline. Returns a paginated list of all agent invocations with metrics.
Retrieve the full execution transcript for a session. Only available when session persistence is enabled.
Parameter descriptions missing or incomplete across all tools. claude_code_timeline documents 'offset', 'limit', 'run_id' with minimal context. claude_code_sessions 'status' parameter has description but no enum constraint or valid values listed. create_claude_code_mcp 'permission_mode' accepts free-form string ('bypassPermissions|acceptEdits|default|plan|dontAsk') without schema enum constraint.
No output schemas documented for any tool. Callers cannot know what fields to expect from claude_code (session_id, result_text?), claude_code_report (HTML string? file path?), claude_code_timeline (list of what structure?), or claude_code_transcript (string? structured log?). This forces LLMs to guess downstream parameter usage.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | F | 45 | 2026-07-28+ | v2 |
Factory tool to dynamically create new specialized Claude Code MCP servers. Only available when CLAUDE_FACTORY_ONLY=true.
Pagination parameters (offset, limit) in claude_code_timeline lack documentation of expected ranges, defaults, or total_count return field. Rubric baseline specifies tools returning lists must document page/offset, limit, and total count to avoid context window exhaustion. No guidance on max results per call.
No error recovery guidance. Tool descriptions state what they do but not what to do if they fail. E.g., claude_code_reply doesn't explain 'if session_id not found, call claude_code_sessions() first'. Irreversible tools (claude_code, claude_code_reply, create_claude_code_mcp) don't mention confirmation or dry-run support.
Parameter type constraints under-specified. create_claude_code_mcp accepts 'model' as free-form string (users might pass 'claude-3.5-sonnet' vs 'sonnet' vs 'opus'); 'allowed_tools' is comma-separated string with no validation. All params should have explicit enum, pattern, minLength, maxLength in schema.
Tool composition risk: claude_code and claude_code_reply both accept 'prompt' and execute code, but no documentation clarifies their differences or when to use one vs. the other. LLMs will conflate similar tool names (baseline issue from review:name-clarity).
Descriptions do not explain dependencies between tools. E.g., claude_code_reply requires a valid session_id from a prior claude_code call, but the description doesn't guide LLMs to understand this prerequisite. claude_code_transcript is only available 'when session persistence is enabled', but that condition isn't documented in the tool description.