CLI-first Staff Review & Decision Assistant with MCP server for staff-level code review and decision making
greybeard exposes 4 tools via MCP with generally clear intent but inconsistent quality. All tools have descriptions and input schemas are present, but parameter descriptions vary in completeness. Tool names follow verb_noun convention (review_decision, self_check, coach_communication, list_packs). Schemas use proper JSON Schema format with types. However, several parameters lack descriptions, output schemas are not documented, and error handling guidance is absent.
Get help communicating a concern, risk, or decision to a specific audience. Returns suggested phrasings that are collaborative rather than blocking.
List all available content packs (built-in and installed).
Review a decision, design document, or git diff from a Staff Engineer perspective. Returns structured markdown with risks, tradeoffs, and questions to answer before proceeding.
Review your own proposal or decision before sharing it. Acts as your internal critic, surfacing weak arguments, unstated assumptions, and questions your reviewer will ask.
Output schemas not documented. Tools describe what they return in prose (e.g., 'Returns structured markdown'), but do not specify the exact JSON structure, field names, or types that clients should expect. LLMs cannot plan downstream operations or extract specific fields without knowing the response schema.
Parameter 'context' in review_decision and self_check marked 'Optional' but lacks description text explaining what context means or how it influences the review. Parameter 'input' in self_check is marked optional with description 'Optional: supporting document or draft', but the required parameter 'context' is not similarly clarified. Inconsistent optional/required semantics.
list_packs has empty properties object (no parameters documented), but description says 'List all available content packs' without specifying what fields the response contains, what a 'pack' object looks like, or whether pagination is supported. Minimal parameter validation.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 64 | 2026-07-28+ | v2 |
| 2026-03-09 | C | 61 | 2024-11-05+ | v1 |
No error handling guidance. Tools do not document what errors can occur, how to recover, or whether errors are retryable. For example, if an invalid pack name is passed, does the tool return a user-friendly message or a raw exception? If the LLM calls review_decision with malformed input, what happens?
Default values ('staff-core', 'mentor-mode') are hardcoded in schemas but not explained. What is 'staff-core'? Why is it the default? What does 'mentor-mode' pack contain? LLMs cannot understand when to override defaults without description.