A local daemon that turns a task into a PR across Claude Code, Codex, Gemini, Qwen, and Droid. Handles stage dispatch, hook-time rules, progressive MCP proxy, and more.
Gobby exposes 9 tools with severe definition quality issues. Most tools lack proper input schemas with complete type and constraint information. Descriptions are present but minimal (10-70 chars), below the 194-char baseline for production tools. Critical issues: (1) Multiple file-editing tools (write_file, edit_file, replace, edit, write, notebook_edit) with overlapping responsibilities and missing parameter documentation, violates single-responsibility and composition patterns. (2) Tool input schemas are incomplete: 'replace' and 'edit_file' lack documented parameters beyond 'file_path' or 'target_file'; actual required parameters (content, start_line, end_line, etc.) are not visible in the schema. (3) No output schemas documented for any tool, LLMs cannot plan downstream operations. (4) 'Skill' tool accepts 'args' as a bare string with no format specification, type constraints, or examples. (5) 'call_tool' accepts 'arguments' as 'string|object' without clarity on when to use each or how to format string variants. (6) No error handling guidance, tools do not document recovery steps, retryability, or what errors mean. (7) No pagination or result limiting for 'list_tools', risking context window explosion. (8) Security concern: 'call_tool' and 'Skill' do not document permission gates or validation rules.
Resolves and executes Gobby-owned skills through local DB or MCP gobby-skills server
Execute a tool with optional pre-validation and workflow enforcement
Edit a file or resource
Edit a file
List tools for a specific server with progressive discovery format
Edit a notebook file
Replace content in a file
Multiple overlapping file-editing tools (write, write_file, edit, edit_file, replace, notebook_edit) violate single-responsibility principle. LLMs cannot reliably distinguish when to use 'edit' vs 'edit_file' vs 'replace', the naming alone does not convey intent.
Input schemas incomplete or missing for 6 of 9 tools. 'replace' and 'edit_file' have no visible schema beyond file_path, actual parameters like content, start_line, end_line are undocumented.
No output schemas documented for any tool. LLMs cannot predict return types or structure, blocking composition of multi-step workflows. Agent cannot know if write_file returns {success: bool, file_id: string} or {written: int} or {error: string}.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 47 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 34 | - | v1 |
Write to a file or resource
Write content to a file
Descriptions are minimal (10 - 35 chars), well below 194-char baseline. 'edit' = 'Edit a file or resource' (25 chars) lacks WHAT data it modifies, WHEN to use it vs write/edit_file, any prerequisites, or whether it's idempotent.
call_tool accepts 'arguments' as either string or object with no guidance on format. Does 'arguments' as string require JSON? YAML? How does the tool distinguish? No examples or constraints provided.
Skill tool: 'args' parameter accepts bare string with no format specification, length limit, or validation rules. LLM cannot determine valid input, is 'args' a JSON string, shell args, or plain text?
list_tools has no pagination or result limiting documented. If a Gobby instance proxies 100+ MCP servers, list_tools could return thousands of tools without limit, exhausting context window. No mention of limit parameter or total count in response.
No error handling guidance across any tool. Tools do not document: (1) retryability (is a write_file timeout retryable?), (2) recovery steps (what to do if write_file fails due to permission denied?), (3) error categories (user error vs system error vs transient).
No permission gates or scope declarations. call_tool and write_file do not document what authorization is required. An agent could theoretically call any tool on any server or write to any path. No indication of least-privilege constraints.
'call_tool' proxy design with optional 'enforce_workflow' and 'strip_unknown' flags creates hidden behavior. LLM cannot know what these flags do without source code inspection. Should these be enum choices? Should enforcement be mandatory? Undocumented behavior invites misuse.
Generic names: 'edit', 'write' are too vague. Both are write operations but unclear how they differ. Per naming rubric, overlapping names force LLM reasoning and errors. Recommend: consolidate into single 'write_file' or 'update_file' tool with explicit mode parameter.