CLI for CodeForge development workflows with codebase indexing, container management, and Claude Code integration
CodeForge is an MCP server that wraps a CLI tool for managing Claude Code sessions, tasks, plugins, configuration, code indexing, and devcontainers. The tool definitions are inferred from CLI command registrations rather than explicitly defined in MCP-native format. While 30 tools are listed, critical quality issues pervade: (1) descriptions range from adequate to vague but lack LLM-optimized guidance on when to use each tool vs. alternatives; (2) input schemas are present but minimalist, most parameters lack type constraints, enums, ranges, or dependency documentation; (3) output schemas are entirely undocumented, LLMs cannot predict what fields will be returned or how to chain tools; (4) no error handling guidance, tools fail silently or with generic messages; (5) no indication of side effects, idempotency, or recovery paths; (6) composition is poor, related tools (e.g., session list/search/show) lack clear disambiguation; (7) tool registration appears to be implicit via command handler registration rather than explicit MCP tool definitions with full JSONSchema. Source: CLI command files like cli/src/commands/session/search.ts show minimal metadata. The server conflates a CLI wrapper with MCP tool design patterns.
Deploy configuration files from workspace to system
Show current Claude Code configuration
Stop a running CodeForge devcontainer
Execute a command inside a running devcontainer
List running CodeForge devcontainers
Rebuild a CodeForge devcontainer
Open an interactive shell in a running devcontainer
Start a CodeForge devcontainer
Build or incrementally update the codebase symbol index
No documented output schemas. LLMs cannot predict response structure or field names. Tools like 'session search', 'task list', 'index search' lack documented returns, forcing LLMs to guess what fields are available for downstream chaining.
Minimal parameter descriptions lacking actionable guidance. Most parameters have bare descriptions ('Session identifier', 'Search query') without constraints, formats, or examples. Missing: min/max for numeric params, enum values for constrained inputs, regex patterns, required vs. optional distinction.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 36 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 23 | - | v1 |
Remove the codebase index database
Search for symbols in the codebase index
Show all symbols in a specific file
Show codebase index statistics
Show codebase directory tree with symbol counts
Search and manage plans
Show Claude Code plugin agents
Disable a Claude Code plugin
Enable a Claude Code plugin
Manage Claude Code plugin hooks
List Claude Code plugins
Show details of a Claude Code plugin
Show Claude Code plugin skills
Proxy commands into a running container
List Claude Code sessions
Search Claude Code session history
Show details of a Claude Code session
Show token usage for Claude Code sessions
List and manage tasks
Search and manage tasks
Show details of a specific task
No error handling guidance. Tools like 'container exec', 'proxy', 'config apply' modify state or execute commands but provide no indication of failure modes, recovery steps, or what the LLM should do on error. No actionable error messages or classification (retryable vs. fatal).
Destructive operations lack confirmation or dry-run support. 'index clean' (removes database) and 'config apply' lack dry-run or confirmation prompts. Agents can irreversibly delete data without warning.
Poor tool composition and disambiguation. Multiple tools operate on the same resource (e.g., session list, session search, session show) with no clear guidance on when to use each. 'session search' vs 'session list', when should the LLM choose one? Output schemas not documented, so chaining is speculative.
Tool definitions appear inferred from CLI registration rather than explicitly defined in MCP format. Source code shows command handler registrations but no explicit MCP tool_definition messages with full JSONSchema. This may indicate tool metadata is auto-generated or missing.
Missing idempotency and side-effect documentation. Tools like 'plugin enable', 'plugin disable', 'container up', 'config apply' modify state but do not declare whether they are idempotent (safe to retry) or what side effects occur. LLMs cannot safely retry on ambiguous failures.
Generic tool names lack specificity. 'proxy' is vague, what does it proxy? Does it forward commands to a container? The name alone does not convey intent. Should be 'container_proxy_command' or 'exec_in_container'.
No pagination or limit guidance. Tools like 'session list', 'task list', 'plugin list', 'index search' return lists but descriptions do not specify pagination, default limit, or max result size. LLMs may request 1000 items, blowing context.
Tool annotations missing. No hint of which tools are read-only (safe, no side effects), which are destructive, or which require confirmation. Helps LLM risk-assess before invoking.