Multi-AI consensus tool: MCP server that queries multiple AI models in parallel, synthesizes responses for better accuracy, and reduces AI bias through ensemble decision-making
Single tool with adequate schema and description, but moderate issues around error handling and input validation guidance. The tool `ai_council` has a clear purpose and documented parameters, but descriptions lack specificity about when/why to use it vs alternatives, error recovery paths are generic, and parameter constraints are not well-communicated. The schema is present and properly typed, placing this in the fair-to-good range, but gaps in error guidance and composition patterns prevent a higher score.
A tool that consults multiple AI models in parallel, then uses one of them to synthesize the results into a single, high-quality answer. Use this for complex questions requiring deep analysis and verification.
Parameter descriptions lack format/constraint guidance. 'context' and 'question' are described at a high level but do not specify length limits, expected structure, or examples of what makes good context/question vs poor input.
Error handling does not guide LLM toward recovery. ErrorResponse returns generic error codes (UNKNOWN_TOOL, INTERNAL_ERROR) with minimal context. When synthesis fails or model responses are empty, the LLM receives no hint about what caused failure or what to try next.
Tool description does not explain when to use this tool vs direct single-model calls. 'Complex questions requiring deep analysis and verification' is vague. Missing guidance: minimum complexity threshold, when consensus adds value, acceptable latency/cost tradeoffs.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 56 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 46 | - | v1 |
Output schema documented in code but not machine-readable from tool registration. The tool description does not state what structure the response will have (success/error union, fields returned, consensus metadata). LLMs must reverse-engineer output structure from example usage.
Input validation is minimal. Code checks for empty strings and basic length (context ≤10k, question ≤5k) but these constraints are not documented in the tool description. LLMs cannot see these limits and may pass invalid input without guidance.
No dry-run or confirmation pattern for irreversible operations. This is a query tool (read-only per risk annotation), but the parallel querying of multiple external APIs consumes credits/quota. No option for LLM to preview cost or request confirmation before execution.