Multi-MCP presents a collection of conversational and analysis tools with partial schema coverage and inconsistent description quality. While tool names are clear (codereview, chat, compare, debate), descriptions vary in depth but generally lack the operational clarity needed for optimal agent planning. The server exposes 6 tools via fastmcp (STDIO-only), all with READ_ONLY risk profiles. Parameter schemas are present but lack critical constraints (enums, min/max bounds). Output schemas are not documented. No tool annotations (readOnlyHint, destructiveHint, idempotentHint) are present. Error handling and recovery guidance are not visible in tool definitions. The codebase shows engineering discipline (80%+ coverage, linting, type checking) but the tools themselves need refinement to meet production-grade agent tool standards.
General chat with AI assistant. Supports multi-turn conversations with project context and file inclusion.
Systematic code review using external models. Covers quality, security, performance, and architecture.
Compare responses from multiple AI models. Runs the same content against all specified models in parallel. Supports multi-turn conversations with project context and file inclusion.
Multi-model debate: Step 1 (independent answers) + Step 2 (debate/critique). Each model provides independent answer, then reviews all responses and votes.
List available AI models. Returns model names, aliases, provider, and configuration.
Get server version, configuration details, and list of available tools.
No output schemas documented for any tool. LLMs cannot plan downstream calls or extract fields without knowing what each tool returns. Current tools lack documented response structures.
Parameter 'step_number' and 'next_action' appear in 4 tools (codereview, chat, compare, debate) without clear semantics. Are they stateful internal IDs? Can an agent safely set them? Should they be auto-managed by the server? Undocumented dependencies violate parameter relationship clarity.
Parameter 'models' (array of strings) lacks enum or constraint documentation. What models are valid? Can an agent pass any string? Should match against output of 'models' tool. Undocumented dependency forces discovery call.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 48 | <=2025-11-25 | v2 |
| 2026-03-09 | D | 56 | - | v1 |
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) despite all tools being marked READ_ONLY risk. Tool annotations were added to MCP spec for LLM planning optimization. Current implementation relies on risk labels in documentation, not structured metadata.
Utility tools 'version' and 'models' have weak descriptions (50 chars each: 'Get server version...' and 'List available AI models...'). Too generic to guide LLM selection. Should explain when to call them (on startup? before tool invocation?) and what they enable.
Parameter 'base_path' (string, description: 'Base path for project context') is under-specified. Is it a filesystem path? Must it exist? Validation behavior on invalid paths? Required or optional? Should have min/max length, pattern, or enum.
Parameter 'thread_id' appears in 4 tools without clear ownership semantics. Is it a session ID the agent should persist? Auto-generated by the server? If state-dependent, violates MCP stateless principle. Needs clarification in descriptions.
No validation constraints visible for numeric parameters. 'step_number' is an integer with no min/max. Can an agent pass step_number=999999? Should be bounded to documented process length.
No error handling guidance documented. What happens if base_path doesn't exist? If a model is unavailable? If thread_id is invalid? Tool descriptions lack recovery paths (e.g., 'If model X is not available, try search_models() first').