Multi-round AI model deliberation server that orchestrates consensus-building debates between multiple AI models with tool-based evidence gathering. Exposes deliberation, decision graph querying, and quality metrics via MCP.
This MCP server has critical gaps in schema completeness, parameter descriptions, and output documentation. Of 9 tools, 6 have either missing or incomplete input schemas, and most lack clear output schema definitions. The 'participants' parameter in 'deliberate' uses an empty oneOf array (invalid schema), indicating incomplete implementation. Descriptions are present but generic (averaging ~100-150 chars), lacking LLM-optimized guidance on when to use each tool or what to expect in responses. No tool includes output schema documentation. Error handling is absent, tools provide no recovery guidance. Security considerations for command execution ('run_command') and file access are not documented. The server reads like an internal prototype rather than a production-ready agent tool suite.
Initiate deliberative consensus where AI models debate across multiple rounds. Models see each other's responses and adapt their reasoning. Supports CLI tools (claude, codex, droid, gemini, llamacpp) and HTTP services (ollama, lmstudio, openrouter, nebius).
Get visual file tree representation of directory structure during deliberation. Used by AI models to understand project layout for evidence-based analysis.
Track response quality metrics per model to monitor deliberation performance and model reliability.
List files matching glob pattern during deliberation. Used by AI models to discover codebase structure for evidence gathering.
Query the decision graph memory to retrieve stored deliberation results and consensus decisions (when enabled in config).
Read file contents during deliberation. Used by AI models to access codebase files for evidence-based analysis.
deliberate tool has invalid schema: 'participants' array uses empty oneOf (no valid options defined). This will cause schema validation failures and LLM cannot determine valid participant values.
No output schemas documented for any of the 9 tools. LLMs cannot plan downstream calls or understand what fields to extract. This violates pattern:tool and pattern:response-shaper fundamentally.
run_command tool lacks security documentation and no mention of whitelist enforcement. Description says 'restricted to safe, read-only operations' but no schema shows how restrictions are enforced or what happens if an agent passes an unsafe command.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 47 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 31 | - | v1 |
Execute read-only commands during deliberation. Used by AI models to run shell commands for gathering evidence (restricted to safe, read-only operations).
Search codebase with regex patterns during deliberation. Used by AI models to find code matching specific patterns for evidence gathering.
Set or clear session-scoped model overrides by adapter to customize which models are used for each adapter during deliberation.
query_decisions tool description is vague: 'retrieve stored deliberation results and consensus decisions', when should this be called? What is the expected response structure? Does it require prior calls to deliberate()? No dependency hints.
set_session_models tool has no description of what 'session-scoped' means or how it interacts with the deliberate tool. Input schema uses additionalProperties with string|null but no examples of valid adapter names or model overrides.
read_file and search_code lack parameter descriptions for security constraints. 'path' param in read_file has no mention of exclusion patterns, symlink handling, or max file size limits. LLMs will not infer these constraints.
No error handling documented. What happens if deliberate times out? If a participant model is unavailable? If a file is too large to read? How should an LLM recover? No recovery guides in any tool description.
get_quality_metrics has optional 'model' parameter with no enum or description. What model names are valid? Can it be omitted to get all models? Unclear.
list_files and get_file_tree descriptions do not explain what happens on permission errors or if path does not exist. No guidance on pagination for large directories.