AI agent backend for project management, task tracking, GitHub repository analysis, and team collaboration with WebSocket support
CommitFlow exposes 8 tools with detailed, well-written descriptions that clearly state WHEN to use each tool and what it returns. All tools follow a consistent verb_noun naming pattern (getRepos, getContributors, getProjects, etc.). However, critical gaps exist: NO input schemas are visible in the source code provided, only tool names, descriptions, and parameter lists are shown. The parameter objects (e.g. {"type":"object","properties":{...}}) appear to be schema fragments but are NOT explicitly tied to MCP tool registration or validation logic in the backend code. This means while descriptions are strong (80+ baseline), the actual JSON Schema validation and machine-readable type definitions cannot be verified. Additionally, no output schemas are documented, no error handling patterns are shown, and the tools lack security annotations (readOnlyHint, destructiveHint, idempotentHint). All 8 tools are READ_ONLY according to metadata, but this is stated in comments, not in structured tool annotations. The server is HTTP-based (NestJS backend) but the MCP server implementation itself is not shown, only the tool definitions file and Docker setup are visible.
Retrieve all tasks (filtered by projectId if provided). Results should be detailed and actionable: include id, title, description snippet, status, priority, assignee (if any), createdAt, updatedAt, dueDate, age (days open), linked PR/issue IDs, dependencies, blocker flag, commentsCount, and suggested next action (e.g., 'assign', 'review PR', 'change priority'). Use this when the user requests a list of tasks without a status filter.
Use this function when the user asks about contributors, commit counts, top contributors, or developer activity for a specific repository. This function requires the repository name. The result should provide contributor-level metrics and insights: total commits, commits in the last 30/90 days, PRs opened/merged, issues opened, lines added/removed (if available), recent activity timestamp, and a short note identifying top contributors and potential areas for recognition or triage.
Retrieve tasks with the 'inprogress' status. Provide details including blockers, time-in-progress (age), PR links, assignees, and recommended next steps to complete (e.g., 'needs QA', 'awaiting review'). Can be filtered by projectId.
Retrieve the list of team members along with their task statistics and workload insights. Response should include tasks assigned, tasks by status, overdue counts, recent activity, and a short workload recommendation (e.g., 'overloaded — reassign', 'underutilized — assign new tasks'). Use this when the user asks about assignees, workload distribution.
NO OBSERVABLE INPUT SCHEMAS, Schema objects are shown as JSON fragments in the tool definitions, but the actual MCP server code that registers these tools with validated JSON Schema is not visible. Cannot verify that these schemas are actually enforced, that properties have types beyond 'object', or that required arrays are correctly declared.
NO OUTPUT SCHEMAS DOCUMENTED, Descriptions mention what data will be returned (e.g. 'Result should include contributor-level metrics...') but no formal output schema is declared. LLMs cannot plan downstream calls or extract fields reliably without knowing the response structure. Agents will struggle to chain tools.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 54 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 42 | - | v1 |
Retrieve the list of active projects along with task statistics (todo, inprogress, qa, deploy, done). The response should include per-project insights: task breakdown by status, completion rate (%), overdue task count, blocked tasks count, recent activity, risk level (low/medium/high) and recommended next steps (e.g., 'reassign overdue tasks', 'prioritize critical bugfixes'). Use this function when the user asks about project lists, project details, or which project contains certain tasks.
Retrieve tasks currently in the 'qa' status. These tasks are awaiting quality assurance validation. Include testing readiness, linked PRs/builds, detected issues, blockers, and how long the task has been in QA. The response should help decide whether the task can be approved for deploy, needs fixes, or is blocked. Use this when the user wants to review QA workload, bottlenecks, or release readiness. Can be filtered by projectId.
Retrieve the list of GitHub repositories. Use this function when the user asks about repository names, available repositories, or repo options. The result should include actionable insights for each repo: recent activity (last commit date), open issues/PR counts, active contributors count, primary language, CI status (if available), and a short recommendation (e.g., 'needs PR review', 'archived candidate').
Retrieve tasks with the 'todo' status. Include the same detailed fields as getAllTasks plus suggestions for prioritization (e.g., 'start next week', 'urgent: reassign'). Can be filtered by projectId.
MISSING TOOL ANNOTATIONS, All 8 tools are READ_ONLY and safe to retry (idempotent), but these properties are not exposed as MCP tool annotations (readOnlyHint, idempotentHint). Agents cannot infer safety properties from the protocol.
NO ERROR HANDLING PATTERNS DEFINED, No evidence of categorized error responses (retryable vs user-fixable vs fatal), no recovery guidance, no actionable error messages. If a repo lookup fails or an invalid projectId is passed, agents will not know what to do next.
PAGINATION NOT ADDRESSED, Tools like getRepos, getContributors, getAllTasks, getTodoTasks, getInProgressTasks, getQaTasks return lists but no pagination parameters (limit, offset, page, cursor) or response structure (total_count, next_cursor) are documented. Large result sets will blow context windows.
PARAMETER TYPE DETAILS MISSING, Schema properties show 'type: "string"' but lack length constraints, format hints, or patterns. E.g., 'repo' parameter is just a string, should specify (1) min/max length, (2) allowed characters, (3) whether to resolve from user input. Without these, LLMs may pass malformed values.
NO CHAINING IDS DOCUMENTED, Descriptions say results will include data like 'linked PR/issue IDs', 'linked PRs/builds', etc., but do NOT specify field names (pr_id, issue_id, build_id?). If a downstream tool needs these IDs, agents cannot reliably map output fields to input parameters without extra discovery.