MCP Server using Google AI SDK and repomix for planning and review.
The server defines 5 tools with reasonable naming conventions (verb_noun format: get_next_task, mark_task_complete, plan_feature, review_changes, adjust_plan). However, there are significant gaps in parameter descriptions, output schema documentation, and error handling guidance. All tools have descriptions (satisfies the critical requirement), but most are brief and lack context about when to use each tool or what the LLM should expect in responses. Parameter descriptions are present but generic, they state what the parameter is (UUID, string) without explaining the semantic meaning or contextual usage. Output schemas are not documented anywhere in the visible code; the code returns text responses without structured field documentation. Error handling exists but returns unstructured text without recovery guidance. The tools form a logical chain (plan_feature → get_next_task → mark_task_complete → review_changes → adjust_plan), but the composition lacks clear handoff points and required fields for chaining.
Adjusts an existing feature plan based on user feedback or requirements changes.
Retrieves the next pending task for a given feature. Returns the first pending task with details including effort and parent task information if applicable.
Marks a task as complete for a given feature.
Plans a feature by analyzing a project directory using repomix and generating tasks based on the feature description.
Reviews changes made during feature implementation by analyzing git diffs and generating a review report.
No output schemas documented for any tool. Code returns unstructured text responses without specifying what fields, formats, or structures the LLM should expect. LLMs cannot plan downstream tool calls or reliably extract field values.
Parameter descriptions are generic type annotations (UUID, string) without semantic context. E.g., 'Valid feature ID (UUID) is required' restates the type; it does not explain what a feature represents, how to obtain a valid feature ID, or what happens if an invalid ID is passed.
Error handling returns unstructured text (error.message) without recovery guidance. E.g., 'No tasks found for feature ID X' is a fact, not a recovery hint. LLM does not know: is this retryable? Should I call plan_feature first? Is the feature_id wrong?
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 46 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Tool composition lacks handoff clarity. plan_feature returns no documented feature_id; subsequent tools require featureId parameter. LLM cannot reliably extract the correct ID from the text response.
plan_feature and adjust_plan perform multiple operations in a single tool (analyze repo AND generate tasks; analyze feedback AND modify plan). These should be split for clarity and composability.
No parameter constraints for free-form strings (feature_description, adjustment_request). adjustment_request only has minLength=1; no maximum length, format guidance, or examples. Invites invalid input.
No idempotency guidance. mark_task_complete and adjust_plan have write side effects, but lack documentation on whether they are safe to retry or could cause duplicate effects.