Multi-model debate orchestrator for getting opinions and reviews on code and tasks. Supports OpenAI, Gemini, and Anthropic models with intelligent token-based model selection and multi-round debate capabilities.
The server defines 2 tools with reasonable descriptions and input schemas. Both tools follow a verb_noun pattern (sage-opinion, sage-review) and have complete JSON Schema input definitions using Zod. However, there are significant gaps: (1) descriptions lack actionable context about WHEN to use each tool vs. the other, (2) output schemas are not documented in the tool definitions, (3) no parameter-level validation guidance (e.g., what constitutes valid paths, how many paths are too many), (4) error handling is mentioned in the code but not surfaced in tool descriptions to guide LLM recovery, and (5) the tools accept arrays of paths but provide no pagination or limit constraints on context size despite explicitly mentioning 'Do not worry about context limits', this invites resource exhaustion.
Send a prompt to sage-like model for its opinion on a matter. Include the paths to all relevant files and/or directories that are pertinent to the matter. IMPORTANT: All paths must be absolute paths (e.g., /home/user/project/src), not relative paths. Do not worry about context limits; feel free to include as much as you think is relevant. If you include too much it will error and tell you, and then you can include less. Err on the side of including more context. If the user mentiones "sages" plural, or asks for a debate explicitly, set debate to true.
Send a prompt to sage-like model for a detailed review of code, design, or approach. Include the paths to all relevant files and/or directories that are pertinent to the matter. IMPORTANT: All paths must be absolute paths (e.g., /home/user/project/src), not relative paths. Do not worry about context limits; feel free to include as much as you think is relevant. If you include too much it will error and tell you, and then you can include less. Err on the side of including more context. If the user mentions "sages" plural, or asks for a debate explicitly, set debate to true.
Output schema not documented. Tool descriptions do not explain what the response contains, field names, or structure. LLMs cannot plan downstream actions without knowing the response schema.
Descriptions lack differentiation guidance. Both tools have nearly identical descriptions; it is unclear when to use sage-opinion vs sage-review. An LLM cannot distinguish between them based on the provided text.
No limits or constraints on paths parameter. Description says 'do not worry about context limits' but provides no guidance on maximum array length, file count, or total token budget. This invites agents to submit unbounded context that can exhaust resources or fail silently.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 48 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 21 | - | v1 |
Error handling not documented in tool descriptions. Code shows error paths (e.g., packFiles throws on overload) but tool descriptions do not explain what to do if context is too large, invalid paths are passed, or the model call fails.
No parameter-level validation hints. The paths parameter requires absolute paths per the description, but there is no validation constraint in the schema (e.g., regex pattern or minLength check) and no actionable error message documented for when a relative path is passed.
Debate parameter purpose underspecified. The description says 'set debate to true when user mentions sages plural' but does not explain what debate mode does differently from the default, what outputs change, or how long it takes. LLM may invoke debate mode without understanding consequences.