This GitHub MCP server has 26 well-defined tools with consistent naming patterns and generally adequate schemas. Most tools follow verb_noun naming (list_issues, get_pull_request, create_issue) and include input schemas with typed parameters. However, there are significant gaps: (1) Tool descriptions are minimal (typically 5-15 words), providing limited context about WHEN to use each tool vs. similar alternatives; (2) Parameter descriptions exist but are often trivial ('Repository owner' vs. actionable guidance); (3) Output schemas are completely undocumented, LLMs cannot predict what fields to expect from responses; (4) Error handling provides no recovery guidance; (5) No tool annotations (readOnlyHint, destructiveHint, idempotentHint) despite obvious WRITE vs. READ distinctions. The server successfully avoids major antipatterns (no hardcoded examples in descriptions, proper enum constraints for state/sort fields, required parameters clearly marked). However, it falls short of production-grade quality due to missing output documentation and thin descriptions that leave LLM tool selection suboptimal.
Output schemas completely undocumented. LLMs cannot predict response structure, forcing them to reason through returned JSON without guidance. This increases token consumption (LLM must parse and extract fields) and raises hallucination risk when downstream tools depend on specific response fields.
Document output schemas for every tool. Specify which fields are returned, their types, and which ones are safe to chain into downstream tools. Example: 'Returns: {id (number), title (string), state (string: open|closed), created_at (ISO 8601 date)}'. This enables LLMs to plan tool composition without hallucinating field names.
Expand tool descriptions to 80-150 chars, answering: WHAT does it do? WHEN to use it instead of similar tools? WHAT does it return? Example: 'List all issues in a repo, filtered by state/labels. Use this for discovery; use search_issues() for complex queries. Returns paginated list of issues (max 100 per page).'
Add tool annotations. Mark all WRITE operations with 'destructiveHint': true (create_issue, update_issue, merge_pull_request, create_or_update_file, etc.) and all READ operations with 'readOnlyHint': true. Mark idempotent operations like mark_notifications_read with 'idempotentHint': true.
Document pagination behavior explicitly. For list_* tools, add to description: 'Results paginated to max 100 per page (default 30). Accepts per_page (1-100) and page parameters. Returns total_count for result estimation.'
Enhance error messages with recovery hints. Instead of returning raw 'Error: Not found', return actionable guidance: 'Issue #999 not found in repo. Check issue_number is correct or use list_issues() to find valid IDs.'
Provide parameter validation constraints in descriptions. E.g., for 'merge_method', describe: 'Must be one of: merge (create merge commit), squash (squash commits), rebase (rebase onto base branch). Default: merge.' This prevents LLMs from passing invalid values.
Score history
Overall score trend
↓ 2 points across a rubric change (v1 → v2)
49/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
F
49
2026-07-28+
v2
2026-03-09
D
51
-
v1
get_pull_request
read onlyauthsource verified65/100
Get a specific pull request
get_pull_request_filesread onlyauth50/100
Get the list of files changed in a pull request
get_repositoryread onlyauthsource verified63/100
Get repository information
get_workflow_runread onlyauth50/100
Get details of a specific workflow run including jobs
Tool descriptions are extremely minimal (5-15 words). LLMs cannot disambiguate between similar tools (e.g., 'List issues' vs 'Search issues') or understand when each should be used. Production baseline is 50-200 chars with context about WHEN and WHY to call the tool. Current avg ~34 chars is well below baseline.
No pagination limits documented. list_issues, list_pull_requests, list_repositories, list_workflows accept per_page param but provide no guidance on maximum values or default behavior. Without explicit caps (e.g., 'per_page max 100, default 30'), LLMs may request 1000+ items, wasting context window.
Error handling provides no recovery guidance. When a tool fails (e.g., 'Error: Not found'), LLM receives no actionable next step. Should include hints like 'User not found. Try search_users() with partial name' or 'Repository not found. Available repos: ...'
Parameter descriptions are trivial and non-actionable. E.g., 'Repository owner' provides no guidance on format (is it a username, email, GitHub ID?). Production baseline includes format constraints, ranges, and examples in descriptions, not just a 1-2 word restatement of the parameter name.
Add examples to descriptions showing common use cases. E.g., 'search_issues': 'Search for issues using GitHub query syntax. Example: query="is:open label:bug" returns all open bug reports. Use this after list_issues() if you need advanced filtering.'
Consider wrapping common multi-step workflows into composite tools. E.g., create_issue → add_issue_comment → add_labels could be a single 'create_issue_with_details' tool to reduce round-trips and token consumption.
Return chaining IDs in all responses. If get_issue() returns issue data, ensure it also returns owner and repo (needed for add_issue_comment). If create_pull_request() succeeds, return pull_number, owner, repo, and branch refs so agent can immediately call create_review().
Cap result limits aggressively. Document: 'list_issues returns max 50 issues per page (not 100) to preserve context window. For bulk operations, use pagination rather than requesting all results.'
Add a 'context' parameter to get_* tools to control verbosity. E.g., get_issue(issue_number, context='minimal') returns {id, state, title} vs context='full' returns everything. This lets LLMs control token consumption.
Clarify parameter mutuality. E.g., for list_issues, document: 'Use either 'labels' (comma-separated) OR state enum, not both' if that constraint exists, or clarify that they compose with AND logic.
Add readOnlyHint annotation to all READ operations (get_*, list_*, search_*) and destructiveHint:true to WRITE operations. This enables LLMs to reason about safety and retry logic without reading descriptions.