Multi-model MCP server — compare, vote, and synthesize across GPT, Gemini, Claude, and local models from one terminal
HydraMCP demonstrates strong pattern knowledge and well-crafted descriptions for complex multi-model reasoning tools. Tool names are verb-based and clear (ask_model, compare_models, consensus). However, input schemas are not visible in the provided code, only Zod schema imports are referenced (e.g., 'askModelSchema.shape'). This creates a critical gap: per the hard rules, schema score must be 0 when schemas are not visible in source. Additionally, tool registration uses only schema.shape (a Zod construct) without explicitly exposing JSON Schema, making it impossible to verify parameter types, required fields, and constraints for downstream clients. Descriptions are exceptionally long and detailed (often 400+ chars), which is good for LLM understanding but violates the rubric baseline of 10 - 1024 chars; the extreme verbosity (e.g., ask_model at ~750 chars) risks token waste. Error handling guidance is present in descriptions but not in structured response schemas. Output format is documented in descriptions ('Markdown with...'), but structured output schema definition is not visible. Naming is excellent across all tools, verb_noun convention with clear semantics. Parameter descriptions are embedded in Zod schema files (not visible here), so parameter annotation quality cannot be verified.
Offload file analysis to a worker model. The file is read server-side — it never enters your context window. You send a file path and a question, and get back only the analysis. OUTPUT: Markdown with the model's analysis. If max_response_tokens is set and compression occurred, includes distillation metadata.
Query any AI model with a prompt. Returns the model's response with metadata. OUTPUT: Markdown with the model's response, latency, and token usage. If max_response_tokens is set and compression occurred, includes distillation metadata (original tokens, compressed tokens, compressor model, compressor latency). Shows "Saved: X tokens (Y% smaller)" when compression is active. Shows "(cached)" when response is served from cache. WHEN TO USE: When you need another model's perspective, analysis, or capabilities. Set max_response_tokens to control how much of your context window this response consumes — the response will be distilled by a fast model to fit the budget while preserving code, file paths, errors, and actionable details. Set include_raw=true to see both compressed and original responses for quality verification. FAILURE MODES: - "Model query failed (4xx/5xx)" → The model or provider is unavailable. Try a different model or check that CLIProxyAPI/Ollama is running. - "circuit breaker open" → The model failed too many times recently. Try a different model or wait for automatic recovery. - Compression silently skipped → If the compressor model is unavailable or the response already fits the budget, the raw response is returned unchanged. This is not an error.
Query 2-5 models in parallel with the same prompt. Returns side-by-side comparison with latency and token metrics.
Input schemas not visible in source code; only Zod schema imports referenced (e.g., askModelSchema.shape). This prevents verification of parameter types, required fields, enums, and constraints for downstream MCP clients.
Tool descriptions exceed baseline length (baseline 10 - 1024 chars; most A+ tools 50 - 200 chars). ask_model description is ~750 chars, session_recap ~800 chars. Extreme verbosity wastes tokens and risks context window dilution. Compress to 200 chars core description + move detailed failure modes and use cases to parameter annotations or separate docs.
Two tools (consensus, session_recap) are inferred only from package.json entries and import statements, not explicitly visible in src/server.ts server.tool() registration. Verify explicit registration for all 8 tools in provided code.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-21 | F | 42 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 43 | - | v1 |
Query 3-7 models and aggregate responses using voting strategy (majority/supermajority/unanimous). Returns consensus answer with confidence score.
List all available models across all providers. Run this first to see what you can query.
Read previous Claude Code sessions from disk and generate a smart-sized recap using a large-context model. Claude never sees the raw session data — only the distilled summary. OUTPUT: Returns markdown starting with "## Session Recap" containing sections: Project State, What Was Built, Key Decisions, Errors Resolved, Unfinished/In Progress, File Map. Empty sections are omitted. Output size is auto-calculated (1K-30K tokens) based on session density. WHEN TO USE: At the start of a new session when the user asks to restore context, recall previous work, or continue where they left off. FAILURE MODES: - "No recent project detected" + list of available projects → Retry with an explicit project path from the list. - "Project directory not found" + available projects → The project path was misspelled or encoded wrong. Retry with a path from the available list. - "No session files found" → The project directory exists but has no sessions. Try a different project. - "No models available" → CLIProxyAPI or Ollama is not running. Tell the user to start their model provider. - "Session Recap Failed" with error details → Both summarization passes failed. Retry with fewer sessions (sessions=1) or a different model. - "Triage Only" heading → Partial success. The triage pass worked but the full recap failed. The output still contains useful structured data. Do not retry.
Read a file with smart caching and token budgets. Returns file content, optionally distilled to fit a max_response_tokens budget.
Query 2-5 models in parallel, then combine their best ideas into one answer. Returns a synthesized response that's better than any single model.
Output schema documentation is narrative (e.g., 'Markdown with the model's response, latency, and token usage') but not formally defined. LLMs cannot infer structured field names, types, or pagination from prose. Provide explicit output schema (JSON Schema format) showing: response text field, latency field (type: number), token usage object, compression metadata object, cache status flag.
Error handling is described narratively in tool descriptions but not structured. Error messages returned in code (e.g., 'Error: <message>') are raw exception strings, not actionable guidance. Pattern:recovery-guide requires: categorize errors (retryable/user-fixable/fatal), suggest next action, include invalid value + constraint. Example: 'Model query failed (4xx/5xx)' should return structured error with: { code: 'model_unavailable', suggestion: 'Try different model or verify CLIProxyAPI running', retryable: true }'.
Parameter descriptions are embedded in Zod schema files (not provided). Cannot verify that all parameters have descriptions, type constraints, enums, ranges, or format specs. Rubric baseline: 100% of A+ tool params have descriptions; parameters without min/max constraints risk LLMs passing absurd values (e.g., max_response_tokens=999999).
Responses use generic text/plain MCP content blocks. No structured output with typed fields (e.g., content: [{type: 'text', text: '...'}], metadata: {latency_ms: 1234, tokens_saved: 50}). Baseline pattern:response-shaper requires returning structured objects so agents can extract fields without parsing prose.
No tool annotations visible (readOnlyHint, destructiveHint, idempotentHint per MCP spec 2026-07-28). All tools are READ_ONLY and should declare { readOnlyHint: true } in input schema. No side effects noted, but if any tool modifies state (e.g., caching behavior), mark appropriately.