An autonomous MCP-powered agent that automates GitHub contributions end-to-end, from issue analysis and code changes to pull requests, conflict resolution, and review updates.
GitPilot MCP has 15 tools with basic registration via @mcp.tool() decorators. All tools have descriptions and input parameters with types, but there are significant gaps: (1) Output schemas are not documented for any tool, responses are returned as dicts/strings without schema declarations. (2) Many parameter descriptions lack actionable detail about constraints, valid ranges, or expected formats. (3) Error handling is minimal, most tools raise generic exceptions without recovery guidance. (4) Parameters accept free-form strings where enums would be better (e.g., branch_name, base branch names). (5) Tool descriptions average ~80 chars, below the 194 char baseline for A+ tools. (6) Several tools conflate multiple responsibilities (e.g., create_branch creates AND checks out; sync_with_upstream modifies state AND rebases). The server is functional but requires substantial refinement for production LLM usage.
No output schemas documented for any tool. Responses are returned as untyped dicts/strings. LLMs cannot reliably plan downstream tool calls or extract required fields without knowing response structure.
Parameter descriptions lack actionable constraints. E.g., 'repo_input' accepts 'owner/repo format or full URL' but no validation rules or examples of failure modes are documented. Parameters named 'base' and 'branch_name' should be enums or include explicit allowed values.
Add formal output schemas to every tool. Document return type as a JSON Schema object. Example for get_issue: {"type": "object", "properties": {"number": {"type": "integer"}, "title": {"type": "string"}, "body": {"type": "string"}, "state": {"type": "string", "enum": ["open", "closed"]}, "labels": {"type": "array", "items": {"type": "string"}}}}
Expand tool descriptions to 150-250 chars. Include WHAT (purpose), WHEN (use case), and RETURN (what the agent gets). Example: 'Get a GitHub issue. Call this after identifying an issue number in a repository. Returns: issue number, title, body, state, and labels. Use title and body to plan code changes.'
Constrain free-form parameters with enums or regex patterns. Replace base, branch_name, and similar string params with explicit allowed values documented in the description. E.g., base: 'Base branch name (typically main, master, develop)'.
Split create_branch into two tools: create_branch (just creates) and checkout_branch (just checks out). Same for sync_with_upstream → add_upstream_remote + rebase_on_upstream. Single responsibility per tool.
Add error handling that returns structured recovery guidance. E.g., apply_patch failure should return {"success": false, "error": "patch_conflict", "message": "Patch conflicts with current state at lines X-Y. Review the diff and rebase manually, then retry.", "next_steps": ["run_tests", "apply_patch"]}
Convert retry_fix_with_tests output from text instructions to a structured response: {"status": "ready", "max_retries": 3, "retry_count": 0, "expected_workflow": ["run_tests", "apply_patch", "commit_and_push"]}
Score history
Overall score trend
↑ 16 points across a rubric change (v1 → v2)
49/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
F
49
2026-07-28+
v2
2026-03-09
F
33
-
v1
auth
source verified
63/100
Get details about a GitHub issue.
get_repo_contextread onlysource verified65/100
Search the indexed repository for relevant code chunks.
index_repowritesource verified70/100
Index a repository for semantic code search. The index is stored in a centralized location to avoid polluting the repository with .rag folders.
Tool descriptions are too short (avg 80 chars, baseline 194 chars). E.g., 'Clone a repository to the workspace' lacks context on why/when to call vs similar tools, what workspace means, or what the return value enables. Descriptions should include WHAT, WHEN, and RETURN VALUE.
Multi-responsibility tools violate single-concern principle. create_branch both creates AND checks out a branch (see code: repo.git.checkout('-b', branch_name)). sync_with_upstream both creates a remote AND rebases. Split these into separate tools so agents compose them explicitly.
Minimal error handling. Most tools raise generic exceptions (RuntimeError, GithubException) without recovery guidance. E.g., 'Patch failed' provides no hint on what went wrong or what to try next. Error messages should categorize (retryable vs fatal) and suggest recovery actions.
retry_fix_with_tests returns only instructions (text) rather than a structured response. Does not clarify what fields the agent should check, or return state that downstream tools expect. Tool output should be machine-parseable JSON, not prose instructions.
No pagination support. get_repo_context accepts 'k' param but returns only k results without a next_cursor or total_count. If k=6 but there are 100 matching chunks, the agent cannot iterate. Large unbounded result sets risk context window exhaustion.
Security: fork_repo and create_pull_request use gh.get_user() implicitly (inferred from GITHUB_TOKEN) with no explicit permission gates or audit logging. No documentation of required scopes (e.g., 'repo', 'write:pull-requests'). Agents should understand what they are authorized to do.
No idempotency documentation. Operations like commit_and_push, apply_patch, and fork_repo could produce side effects on retry. Agents may retry on ambiguous failures, tools should declare idempotency properties or offer compensation mechanisms.
No result limiting or cap documented for get_repo_context. If a search query matches many code chunks, responses could grow very large. Tool description should state maximum results and recommend pagination strategy.
get_repo_context
Add pagination to get_repo_context. Accept limit (1-20, default 6) and offset or next_cursor. Return {"results": [...], "total": N, "returned": k, "next_cursor": "..."}. Cap individual result sizes to prevent context window exhaustion.
Document required GitHub OAuth scopes for each tool. E.g., fork_repo requires 'repo' scope; create_pull_request requires 'write:pull-requests'. Add to tool description: 'Requires scope: repo'. This enables agents and operators to understand permission requirements.
Add idempotency notes to destructive/write tools. E.g., commit_and_push: 'Idempotent if no new changes are staged. Safe to retry. If already pushed, subsequent calls return the branch name without double-committing.'
Add rate-limit handling. Tools calling GitHub API should catch rate-limit errors and return a clear message: {"success": false, "error": "rate_limited", "message": "GitHub API rate limit exceeded. Retry after 60 seconds."}. This guides agent backoff strategies.
Validate inputs before API calls. E.g., normalize_repo() should raise ValueError with a helpful message if repo_input is malformed. Return: {"success": false, "error": "invalid_repo_format", "message": "Repository must be 'owner/repo' or a GitHub URL.", "example": "pytorch/pytorch"}
Add permission gate checks. Before fork_repo or create_pull_request, verify the agent/user has the necessary GitHub scopes. Return a clear error if not: {"success": false, "error": "insufficient_scope", "required_scope": "write:pull-requests"}
For write operations, consider adding a dry_run parameter (boolean, default false). E.g., commit_and_push with dry_run=true returns what would be committed without actually pushing. This prevents accidental commits.