AI-powered code review automation for GitLab/GitHub using Claude Code with MCP support
ReviewFlow has 7 tools with complete JSON Schema definitions and descriptions, but exhibits significant quality gaps that prevent higher scoring. All tools have explicit input schemas with type definitions and descriptions present, which is positive. However, descriptions are inconsistent in quality and some are minimal (under 50 chars). The schema for 'add_action' is complex with conditional parameters that are poorly documented, the tool accepts multiple action types (THREAD_RESOLVE, THREAD_REPLY, POST_COMMENT, POST_INLINE_COMMENT) with different required/optional fields depending on type, but this conditional logic is not made explicit. Error handling and output schemas are not documented in the visible code. Tool naming is acceptable (verb_noun pattern for most), but lacks clarity around parameter naming conventions and missing descriptions for return types. The 'record_insight' tool has an ambiguous description ('Use only for genuine recurring review findings') that relies on LLM judgment without clear validation criteria. No tool annotations (readOnlyHint, destructiveHint, idempotentHint) are present despite multiple WRITE operations that could benefit from them.
Add an action to be executed (resolve thread, post comment, reply, inline comment)
Signal that an agent audit is complete
Get all discussion threads from the MR/PR
Get the current workflow state including agents and their status
Record a recurring review finding Ember derived for the answered project into Ember's private per-project memory, so a later answer reuses it without recomputing. Use only for genuine recurring review findings, never arbitrary facts.
Set the current review phase
Signal that an agent audit is starting
add_action tool has conditional parameter logic not documented. Different action types require different combinations of threadId, message, body, filePath, line, but no parameter description clarifies which fields apply to which action type. LLMs cannot infer conditional dependencies.
No output schemas documented for any tool. The visible code shows tool registration but does not show what fields are returned by get_workflow, get_threads, or other tools. LLMs need to know expected response structure to chain tools and extract data.
record_insight description is ambiguous: 'Use only for genuine recurring review findings, never arbitrary facts.' This relies on LLM judgment without clear validation criteria. What distinguishes a 'genuine' finding? How is this enforced?
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 59 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Missing tool annotations. Multiple WRITE tools (start_agent, complete_agent, set_phase, add_action, record_insight) lack destructiveHint or idempotentHint. Agents cannot distinguish which calls are safe to retry without these hints.
No error handling guidance in tool definitions. What happens if jobId does not exist? What should the agent do next? No recovery hints provided.
get_threads description lacks detail: 'Get all discussion threads from the MR/PR', does not specify pagination, expected size, or what fields are returned. LLMs cannot plan downstream tool calls without knowing result structure.
add_action description is vague about the relationship between action type and required parameters. For example, THREAD_RESOLVE likely needs only threadId, but POST_INLINE_COMMENT needs filePath and line. The description does not make these dependencies explicit.