Production-grade MCP server for Wind/Wall/Door multi-perspective debate orchestration
The server has 18 tools with explicit registrations, descriptions, and input schemas. However, quality is inconsistent. Naming is generally good (verb_noun pattern), but descriptions vary widely in quality, some are cryptic (e.g., 'I5:safety override' for force_close_debate) and lack actionable context. Input schemas are present for all tools but lack depth: many parameters lack descriptions or have vague descriptions. Output schemas are not documented anywhere in the provided source. Error handling is not visible in the code sample. The server attempts comprehensive debate orchestration but falls short of production-grade polish expected for a 70+ score. Average tool score: 58/100.
Record turn. role:Wind|Wall|Door. cognition:PATHOS|ETHOS|LOGOS→validates content.
Finalize debate. synthesis:Door's final resolution->closes room.
Consult with a role on a topic within an active debate.
Convene all roles for a multi-perspective discussion on a topic.
Get the system prompt for a debate agent (Wind, Wall, Door).
Extract DecisionRecord from closed debate for decision history.
I5:safety override. reason:logged→force closes any state.
Output schemas not documented. Tools return results but no structured schema documentation visible in source. LLMs cannot plan downstream calls or extract typed fields from responses.
Cryptic descriptions for sensitive operations. 'I5:safety override' (force_close_debate) and 'I4:redact content' (tombstone_turn) reference internal security levels but do not explain what the LLM should do, when to call them, or what happens. Descriptions must be actionable and self-contained.
Parameters lack descriptions or have vague descriptions. E.g., 'tier' in run_debate (no description), 'provider_type' (vague), 'role' in consult (no details on valid values or consequences).
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 56 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 28 | 1.0.0+ | v1 |
State+optional transcript. include_transcript→adds turn history. context_turns:limit depth.
Sync debate turns to GitHub Discussion/Issue comments. Posts new turns as formatted comments with cognition headers. Idempotent: tracks synced turns to avoid duplicates.
Inject human GitHub comment into active debate as context. Fetches comment from GitHub and adds to debate context. Detects which role was replied to for injection typing.
Create room. mode:fixed|mediated. strict_cognition->validate turns.
Mediated mode only. role:Wind|Wall|Door→sets next expected speaker.
Generate ADR from Door synthesis and create PR.
Resolve a question using a debate's decision record.
Resume an auto-orchestrated debate that was paused or interrupted.
Auto-orchestrate debate to completion with role providers.
Search decision records across closed debates.
I4:redact content→hash chain preserved. turn_index:0-based.
Enum constraints not enforced in descriptions. 'role' parameter appears in many tools (Wind|Wall|Door) but is a free-form string. Schema should declare enum with allowed values; description must list them explicitly.
No error handling documentation visible. Tools may fail (e.g., invalid thread_id, permission denied for GitHub tools) but error responses and recovery guidance are not documented. Agents cannot know whether to retry, ask the user, or abort.
Irreversible operations lack confirmation or dry-run support. force_close_debate and ratify_rfc are destructive but no confirmation pattern or dry-run variant is visible.
GitHub token exposure. GITHUB_TOKEN loaded from .env but no evidence that tools sanitize or validate tokens before passing to GitHub API. If responses include partial token strings or error messages expose secrets, LLMs could log or echo them.
Pagination not explicit in output. Tools like search_decisions accept a 'limit' parameter but no documentation of total count, next_cursor, or pagination behavior. Large results could blow context windows.
Tools combining multiple concerns. run_debate and resume_debate orchestrate entire debates, which is appropriate, but github_sync_debate, ratify_rfc, and human_interject blur debate logic with external integrations. Consider separating debate state management from GitHub sync responsibilities.