AI-powered code evaluation and software engineering best practices enforcement. A behavioral MCP that strengthens AI coding assistants via explicit LLM evaluations.
The server exposes only one tool, set_coding_task, with a reasonably detailed schema but significant definition quality gaps. The tool name lacks a clear action verb and the description is vague about state transitions and error handling. While the input schema is present with types and some enums, parameter descriptions are minimal or missing entirely. The tool performs write operations (WRITE risk) but lacks confirmation/dry-run patterns and error recovery guidance. The schema is more verbose than necessary, and critical parameter relationships are not documented. No output schema is visible in the provided code.
Create or update coding task metadata with enhanced workflow management.
Tool name 'set_coding_task' does not start with a clear action verb. 'set' is ambiguous, it could mean create, update, or initialize. Should be 'create_coding_task' or 'update_coding_task' to match LLM expectations.
Tool description 'Create or update coding task metadata with enhanced workflow management' is vague about when to use create vs update mode. It does not explain what 'enhanced workflow management' means, when the tool succeeds/fails, or what the user should do on error.
Parameter 'task_id' is marked REQUIRED when updating, but the schema does not express this conditional requirement. The description says 'REQUIRED when updating existing task' but the JSON schema does not validate this, the parameter appears optional. Conditional parameter requirements must be enforced in the schema or clearly documented with rejection logic.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 44 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Parameter descriptions are missing or generic. 'user_request' has a description, but 'user_requirements' is described only as 'Updates current requirements', vague about format, length, constraints. 'tags' description lacks guidance on valid tag formats, cardinality, or examples. Parameters like 'state' lack descriptions of valid transitions (can CREATED → REJECTED? Can COMPLETED → IN_PROGRESS?).
No output schema is documented. The tool description does not explain what fields the response contains, whether it returns the updated task object, an ID, confirmation, or error details. LLMs cannot plan downstream calls without knowing the response structure.
Destructive write operation (creates/updates tasks) lacks confirmation or dry-run pattern. The tool modifies state but the description does not warn about side effects or offer a preview. No error recovery guidance provided, if the update fails, the tool description does not tell the LLM what to do next.
State parameter has enum values (CREATED, PLANNING, APPROVED, IN_PROGRESS, COMPLETED, REJECTED) but the description does not explain valid state transitions. Can a task go from COMPLETED back to IN_PROGRESS? Can REJECTED tasks be updated? This forces LLMs to guess or fail.
Schema uses 'enum' for task_size but the description says 'defaults to Medium for backward compatibility', the schema does not show a default value. Enum values (xs, s, m, l, xl) are unclear in meaning without documentation of how they map to effort/complexity.
Tool description contains no guidance on error cases. What if task_id does not exist when updating? What if user_request is empty? What if tags exceed a maximum count? No actionable error messages or recovery steps documented.