Multi-LLM council system with peer review and synthesis, exposed as MCP tools and HTTP endpoints
LLM Council defines 3 tools with moderate quality. Tool names follow verb_noun convention (consult_council, verify, health_check). All tools have descriptions ranging from 50-250 chars. However, parameter schemas are partially documented but lack complete type definitions and constraints. The consult_council tool is heavily feature-rich but its schema is complex and not fully validated against JSON Schema standards. The verify tool references Git operations but lacks clear guidance on repository state handling. health_check is minimal but correct. Descriptions are good but could be more concise and specific about when to use each tool versus alternatives.
Consult the LLM Council for guidance on a query. Progress semantics (ADR-046 P3): while deliberating, the tool emits MCP progress notifications — per-model "Stage 1: <model> responded", per-reviewer "Stage 2: <model> reviewed (n/N)", then synthesis progress. Clients that render progress (Claude Code, Cursor) show these live; clients that ignore progress lose nothing (the notifications are best-effort and never affect the result).
Check the health status of the LLM Council service, including OpenRouter API connectivity and configured model availability.
Verify code or documentation changes against a set of criteria. Runs the verification pipeline with confidence tier selection, producing a pass/fail verdict with evidence disposition tracking and optional dissent extraction.
consult_council schema lacks full type definitions and validation rules. Parameters like 'confidence' and 'verdict_type' reference enum values in descriptions but no explicit JSON Schema enum constraint is visible in the source. Parameter 'evidence' is documented as array but inner structure (dicts with keys) is not formally typed.
verify tool parameter 'criteria' is documented as 'List of verification criteria to check' but the expected schema (structure, allowed values, format) is not defined. LLMs cannot infer what constitutes a valid criterion.
Output schemas are not documented for any tool. consult_council returns synthesis results but the structure (fields, types, sample response format) is not explicitly specified. LLMs cannot plan downstream calls or extract required fields.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-21 | C | 60 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 56 | - | v1 |
consult_council description is 250+ characters and includes internal references (ADR-046, ADR-025b, #619) that may not be meaningful to LLMs. The description mixes use case explanation with implementation details (progress notifications, per-model responses). Should be simplified to 50-150 chars with clear WHEN-to-use guidance.
verify tool references Git snapshots and 'repo_root' but does not document whether this is a local filesystem path or remote repository identifier. Also unclear if this tool supports verification of uncommitted changes or only committed snapshots.
No error handling guidance visible. If consult_council fails (API timeout, model unavailable, insufficient budget for confidence tier), what does the LLM see? Are errors retryable? Should the user be prompted?
health_check description states 'Check the health status of the LLM Council service' but does not document what fields are returned or what constitutes a 'healthy' state (e.g., all models available, API connectivity, latency thresholds).