MCP server for AI-powered GitHub project management — 20 tools (169 actions), agent swarm orchestration, GitHub Actions/Releases/Branches, MCP Resources & Prompts, PRD-to-issues pipeline
This server exhibits critical quality gaps across naming, schema completeness, and parameter documentation. While the codebase is extensive with 15 tools, most tools are 'compound' designs that bundle multiple actions under a single tool name. This violates the single-responsibility principle (pattern:tool) and creates ambiguity for LLMs. Descriptions exist but are shallow (often 1 - 2 sentences, 60 - 150 chars vs. the 10 - 1024 production baseline). Schemas are present but incomplete: enums are declared for action types, but parameters lack descriptions in many tools. The server attempts ambitious features (agent orchestration, AI generation, GitHub integration) but the implementation quality does not match the scope. No tool annotations (readOnlyHint, destructiveHint, idempotentHint) are present despite 8 WRITE-risk tools. Output schemas are not documented, and error handling is minimal, no recovery guidance or actionable error messages visible.
Compound tool for agent orchestration and management. Actions: register, list, deregister, check_status.
Compound tool for agent task management. Actions: checkout, release, complete, submit_work_product.
Compound AI tool for analyzing projects, issues, and roadmaps.
Compound AI tool for generating and managing PRDs, tasks, and features. Actions: generate_prd, enhance_prd, parse_prd, add_feature, get_next_task, analyze_complexity, expand_task, create_traceability_matrix.
Compound AI tool for planning sprints, roadmaps, and project strategy.
Tool for managing automation rules and workflows. Actions: create, list, get, update, delete, enable, disable.
Compound tool design violates single-responsibility principle. All 15 tools use action-based routing (enum action parameter) rather than one tool per action. This creates ambiguity for LLMs: 'manage_project' with action='create' is less discoverable than a dedicated 'create_project' tool. LLMs must reason about which compound tool to call AND which action to invoke.
Missing tool annotations for destructive operations. 8 tools (manage_project, manage_milestones, manage_issues, manage_automation, manage_branches, manage_releases, manage_workflows, agent_work, agent_manage) are marked WRITE risk but have no destructiveHint annotation in their definitions. LLMs cannot distinguish safe reads from unsafe mutations without this metadata.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 43 | 2025-06-18+ | v2 |
| 2026-03-09 | D | 53 | 2024-11-05+ | v1 |
Tool for managing branch protection rules and branch operations. Actions: create, list, get, update, delete.
Compound tool for managing GitHub issues. Actions: create, list, get, update, delete.
Tool for managing issue labels. Actions: create, list.
Compound tool for managing GitHub milestones. Actions: create, list, get, update, delete, get_metrics, get_upcoming, get_overdue.
Compound tool for managing GitHub projects. Actions: create (create new project), list (list all projects), get (get specific project), update (update project), delete (delete project).
Tool for managing GitHub releases. Actions: create, list, get, update, delete.
Compound tool for managing sprints. Actions: create, list, plan, get, update, get_metrics.
Tool for managing GitHub Actions workflows. Actions: list, get, trigger, enable, disable.
System tool for server health checks and status. Actions: health_check.
Output schemas are not documented. No tool explicitly documents what fields it returns or their types. This forces LLMs to make assumptions about result structure and breaks tool chaining, if manage_project returns project_id but the next tool expects projectId, the agent fails silently or retries.
Parameter descriptions are sparse or absent. Many parameters lack descriptions or constraints: sprint.goals (array of what?), workProduct (object with what fields?), inputs (object structure for workflows?), trigger/actions (object structure for automation?). LLMs cannot validate or construct correct inputs without clear descriptions.
Ambiguous or vague parameter names. Examples: workflowId could be numeric, UUID, or filename; color in manage_labels lacks format (hex? RGB?); tagName vs name in manage_releases is unclear. Per the rubric, 'color' should be suffixed with type (color_hex). These force LLMs to guess or require discovery calls.
AI tools (ai_generate, ai_analyze, ai_plan) lack detail on actions. ai_analyze and ai_plan have no enum of valid actions, no parameter descriptions beyond 'action'. This makes them nearly unusable, an LLM cannot know what actions are supported without trial-and-error or documentation lookup outside the tool definition.
No error guidance or recovery paths. Tools do not document how LLMs should handle failures (e.g., 'issue not found' should suggest calling manage_issues with action='list' to discover valid IDs). Error responses are likely to be raw HTTP/API errors rather than actionable recovery hints.
Missing pagination and result-limiting documentation. No tool describes max result counts, pagination parameters, or offset/limit behavior. Listing all milestones or issues could return thousands of items, exhausting the context window. The rubric baseline recommends cap at 20-50 items with pagination support.