MCP server exposing PR analysis tools for GitHub pull request analysis, metrics computation, and team summaries with optional Jira and Confluence integration
The server defines 11 tools with varying quality. All tools have descriptions and naming follows verb_noun convention (analyze_pr, get_pr_metrics, list_prs_by_jira_ticket, etc.). However, schema quality is inconsistent: most tools have input schemas visible in the source code with typed parameters and descriptions, but several critical gaps exist. Tool descriptions are present but average ~100-150 characters, which is acceptable but could be more specific about WHEN to use each tool and what distinguishes them from similar tools. Parameter descriptions are generally present but lack details on constraints, ranges, and format requirements. Output schemas are NOT documented in the source code, the code returns PRMetrics, TeamSummary, or dict objects but no formal schema is visible. Error handling is minimal: functions return None or empty dicts on failure rather than providing recovery guidance. Composition is reasonable, tools are single-purpose and reference IDs appear to flow between tools (repo/pr pairs, author names, ticket keys), but missing output documentation makes chaining difficult for LLMs to plan. The server handles sensitive operations (batch_analyze_* tools write metrics) but lacks permission gates or audit logging.
Fetch a GitHub PR, run the analysis pipeline (test quality, LLM coverage estimate when AI is on, optional Jira), persist metrics, and return a JSON summary.
Find PR(s) for the Jira ticket, analyze one, and return full metrics + report. Uses GitHub Search for the ticket key, then runs the same pipeline as analyze_pr.
Discover and analyze all merged PRs by an author across an entire GitHub org. Finds PRs via GitHub Search, runs the full analysis pipeline on each, and persists results.
Discover and analyze all merged PRs in a single repository. Fetches merged PRs since since_days ago, runs the full analysis pipeline on each, and persists results.
Return aggregate stats for a single author across all analyzed PRs: PR count, repos, avg coverage, avg quality score, total tests added.
Output schemas not documented. Return types are inferred from code (PRMetrics, TeamSummary, dict) but no formal JSON schema is provided in tool definitions. LLMs cannot plan downstream tool calls or extract required fields without documented output structure.
Error handling lacks recovery guidance. Functions return None or {'error': '<message>'} without suggesting next steps. Per pattern:recovery-guide, errors should tell the LLM: 'Try search_users() with a partial name' or 'Run analyze_pr() first'. Current errors are bare 'No metrics for this PR' without actionable recovery.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 57 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 26 | - | v1 |
Return a combined testing-quality summary across multiple repositories.
Return a markdown report ready to paste into the PR description. Includes testing quality score, coverage, test ratio, and AI analysis (if available).
Load previously computed metrics for a PR from local storage.
Return aggregate testing-quality stats for a repository: average coverage, average quality score, top contributors, trend over time.
List merged PRs by author in an org (discovery only, no analysis).
List PRs that mention the given Jira ticket (e.g. CLOSE-13348). By default includes **open** and merged PRs. Set merged_only=True for merged-only discovery.
Parameter constraints not fully specified. Parameters like 'limit' and 'since_days' accept integers but lack explicit range constraints (min/max). 'merged_only' and 'run_analysis_if_missing' are booleans but descriptions don't explain default behavior clearly. Enums are not used where they would tighten constraints.
Destructive operations lack confirmation or dry-run support. batch_analyze_author and batch_analyze_repo write metrics to storage without a confirm or dry-run parameter. Per pattern:confirmation-request, agents should be able to preview changes before committing.
Tool descriptions do not clarify WHEN to use each tool or how they differ from similar tools. For example, 'analyze_pr' vs 'analyze_pr_by_jira_ticket', the description doesn't explain that one requires a GitHub PR number directly while the other searches by ticket. This forces LLMs to guess.
No pagination guidance for tools returning lists. 'list_prs_by_jira_ticket' and 'list_prs_by_author' accept 'limit' but descriptions don't specify whether results are paginated, if there's a next_cursor or offset parameter, or whether hitting limit means more results exist. Large result sets could exhaust context windows.
No permission gates or audit logging. batch_analyze_* tools write metrics to persistent storage (per code: 'pipeline.save(metrics)') without checking caller permissions or logging who triggered the analysis. Per pattern:permission-gate and pattern:audit-trail, sensitive operations should verify authority and leave traces.
Sensitive GitHub and Jira credentials not mentioned in code snippet, but if they are injected via environment variables (as they should be per pattern:secret-injection), this should be documented in tool descriptions. No visible evidence of server-side secret injection pattern.