MCP server that lets Claude Code delegate heavy-token tasks to DeepSeek, Kimi, GLM, Qwen, Grok, or any OpenAI-compatible model. Claude orchestrates; the delegate does the heavy lifting. Zero dependencies.
This server demonstrates good definition quality with clear, LLM-optimized descriptions and complete input schemas for all tools. Tool naming follows verb_noun conventions, and descriptions are explicit about when and why to use each tool. However, there are gaps in output schema documentation and parameter constraint detail that prevent a higher score. The server correctly avoids exposing secrets as parameters and provides actionable guidance in descriptions. The main limitation is that output schemas are not formally documented, the tool descriptions mention what is returned (e.g., 'List the configured providers and their models with context windows, output limits, and pricing') but the actual response structure is not schema-defined in the code review.
Alias of delegate (kept for v2 compatibility). Delegate heavy, token-intensive tasks from Claude Code to a cheaper model (DeepSeek, Kimi, GLM, Qwen, Grok, or any configured OpenAI-compatible provider). Use when: analyzing large files (>300 lines), multi-file codebase reviews, generating outputs >200 lines, complex reasoning, math, architecture design, or anytime your response would exceed ~4000 tokens. Claude orchestrates; the delegate does the heavy lifting. Pass `task` (read/write/reason) so routing picks the right model.
Alias of delegate_models (kept for v2 compatibility).
Delegate heavy, token-intensive tasks from Claude Code to a cheaper model (DeepSeek, Kimi, GLM, Qwen, Grok, or any configured OpenAI-compatible provider). Use when: analyzing large files (>300 lines), multi-file codebase reviews, generating outputs >200 lines, complex reasoning, math, architecture design, or anytime your response would exceed ~4000 tokens. Claude orchestrates; the delegate does the heavy lifting. Pass `task` (read/write/reason) so routing picks the right model.
List the configured providers and their models with context windows, output limits, and pricing
Output schemas not formally documented. Tool descriptions mention what is returned (e.g., 'List the configured providers and their models with context windows, output limits, and pricing') but the actual JSON response structure is not defined. LLMs cannot plan downstream calls or validate field names without seeing the schema.
Backwards-compatibility aliases (deepseek, deepseek_models) duplicate tool definitions and create ambiguity. LLMs must decide between 'delegate' and 'deepseek', both do the same thing. Per the rubric, 'Avoid multiple tools that do the same thing differently.' Consider deprecating v2 names in documentation rather than registering them as separate tools.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 69 | 2026-07-28+ | v2 |
Parameter 'files' accepts file paths but does not validate them or document the expected format (absolute vs relative, permissions required, max file size). The code reads files server-side (good), but the parameter description does not explain constraints: 'Absolute file paths to read...' is stated but no mention of max total size, timeout for reading, or what happens if a file is unreadable beyond a terse error in the implementation.
No explicit error recovery guidance. The delegate tool can fail if a file is unreadable, a model is unreachable, or streaming is interrupted. Descriptions do not document retryable vs fatal errors or suggest recovery steps. E.g., 'If the model is unavailable, try a different task routing or provider.' currently missing.
Tool descriptions mention a 'baseline' cost metric (seen in buildFooter call in code) but do not explain what it is or how to use it. LLMs cannot understand pricing without explicit definition of what 'baseline' means (e.g., 'cost relative to GPT-4 at $0.10/1K tokens').