Static source inference · medium confidence · detected: Logging
Deprecated protocol patterns detected
Summary
The Zen MCP Server suffers from severe definition quality issues across most tools. While 16 tools are declared, the source code provided does not contain explicit tool registration with complete schemas. Tool names are descriptive but lack the verb-noun pattern baseline (90% of A+ tools). Critically, no input schemas are visible in the provided code, only tool names and brief descriptions. The base_tool.py excerpt shows infrastructure but not actual parameter definitions or output schemas. Descriptions exist but are generic (50-100 chars, well below the 194-char baseline for production tools) and lack context on WHEN to use each tool, error recovery, or prerequisites. Parameter descriptions are completely absent from visible code. The tools appear to be workflow/reasoning tools (think_deep, consensus, debug) but their actual invocation contracts are not defined in the source. Without seeing server.py's full TOOLS registration dictionary with explicit schemas, this evaluation must treat tool definitions as inferred rather than directly visible, triggering the hard cap of 50 per tool, but the actual score is lower due to multiple critical gaps.
Tools (16)
analyzeread onlyauthsource verified30/100
General-purpose code and content analysis
challengeread onlyauthsource verified34/100
Technical challenge and problem-solving assistance
chatread onlyauthsource verified37/100
Interactive development chat and brainstorming
codereviewread onlyauthsource verified39/100
Comprehensive step-by-step code review workflow with expert analysis
consensusread onlyauthsource verified37/100
Step-by-step consensus workflow with multi-model analysis
debugread onlyauthsource verified38/100
Root cause analysis and debugging assistance
docgenread onlyauthsource verified40/100
Step-by-step documentation generation with complexity analysis
No input schemas visible in provided source code. All tools scored 0 on schema dimension. Without explicit JSON Schema definitions, LLMs cannot determine valid parameter ranges, types, or constraints.
Tool definitions not directly visible in provided source. Only tool names and brief descriptions from metadata are shown; actual registration with schema definitions not included.
Provide complete, visible tool registration with JSON Schema input and output schemas. Include server.py TOOLS dictionary or tools/ file with explicit Pydantic model definitions. Example: class ChatRequest(ToolRequest): content: str = Field(description='The message to discuss with the AI assistant'); models: Optional[List[str]] = Field(default=None, description='List of model IDs to use for multi-model consensus'); max_tokens: int = Field(ge=1, le=10000, default=2000, description='Maximum tokens to generate per response')
Expand tool descriptions to 150 - 250 characters, following the pattern: 'What it does' (What does it do?) → 'When to use it' (When should the LLM call it instead of alternatives?) → 'What it returns' (What structure is returned?). Example for consensus: 'Multi-model consensus workflow that queries 3+ LLM providers on the same task and synthesizes their outputs. Use when high-stakes decisions need diverse perspectives; returns consensus_decision, model_opinions: {model_id: string}, confidence_score: 0.0 - 1.0, dissenting_views: string[].'
Document output schemas for all tools. Add a structured 'Returns' section to each tool's docstring or create output_schema.json files. Example for listmodels: Returns array of {provider: str, model_id: str, display_name: str, capabilities: string[], rate_limit_rpm: int, cost_per_1m_tokens: {input: float, output: float}}. This enables downstream tool chaining.
Add parameter descriptions and constraints to tools that accept inputs. For tools without visible inputs (version, listmodels), confirm they truly accept no parameters; if they do, document them. Example for listmodels if it accepts filtering: provider: Optional[str] = Field(None, description='Filter by provider (openai|google|anthropic|cohere). If omitted, returns all providers.'); include_pricing: bool = Field(True, description='Whether to include rate limits and pricing info in response.')
Spec posture evidence
Inferred effective spec: 2026-07-28+.
Relies on Logging (deprecated) - log to stderr or use OpenTelemetry
Score history
Overall score trend
↑ 19 points across a rubric change (v1 → v2)
41/100
Scored
Grade
Overall
Spec posture
Rubric
2026-09-22
F
41
2026-07-28+
v2
2026-03-09
F
22
-
v1
listmodelsread onlysource verified40/100
List all available AI models from configured providers
plannerread onlyauthsource verified37/100
Interactive sequential planner using workflow architecture
precommitread onlyauthsource verified37/100
Step-by-step pre-commit validation workflow
refactorread onlyauthsource verified37/100
Code refactoring and improvement suggestions
secauditread onlyauthsource verified39/100
Comprehensive security audit with OWASP Top 10 and compliance coverage
testgenread onlyauthsource verified39/100
Test generation and test case suggestions
thinkdeepread onlyauth37/100
Step-by-step deep thinking workflow with expert analysis
Descriptions are generic and lack actionable context. Tool descriptions average 35 - 50 chars (below 194-char baseline for production tools). None explain WHEN to use the tool vs similar alternatives, prerequisites, or error recovery paths. E.g., 'Interactive development chat and brainstorming' for 'chat' lacks guidance on when to select it over 'consensus' or 'thinkdeep'. State WHAT the tool does, WHEN to use it, and any prerequisites.'
No parameter descriptions visible. Even tools with documented input params (version, listmodels show 'Input: {}') lack descriptions explaining what each parameter controls, valid ranges, or formats. LLMs cannot infer parameter meaning from names alone.'
No output schema documentation. Tool descriptions do not state what fields are returned, type information, or structure. Without documented return schemas, LLMs cannot plan downstream tool chains or extract necessary data. LLMs need to know what fields to expect so they can plan downstream tool calls and extract the right data.'
Error handling strategy not documented. No tool descriptions include guidance on error recovery, retryability, user-fixable vs fatal errors, or what to do if the tool fails. Try search_users() with a partial name." A raw error code or stack trace gives the agent nothing to act on.'
Tool naming lacks consistent verb-noun pattern. Names like 'chat', 'analyze', 'challenge' are generic; baselines show 90% of A+ tools start with action verbs (get, create, list, search, update). 'analyze' could mean code analysis, sentiment analysis, or cost analysis. LLMs rely on the name to infer what a tool does before reading the description.'
Tool composition unclear. Multiple tools appear to overlap (consensus, codereview, debug, thinkdeep all suggest multi-model reasoning workflows). Without clear descriptions of when to use each vs others, LLMs will waste reasoning cycles deciding between similar tools. If search_users and find_users both exist, the LLM wastes reasoning cycles deciding between them.'
consensuscodereviewdebugthinkdeep
Clarify tool composition by adding 'Alternatives' sections to descriptions. Example for consensus: 'Similar tools: thinkdeep (single-model deep reasoning), codereview (task-specific analysis). Use consensus when you need multi-model perspective; use thinkdeep for step-by-step reasoning from one expert model.'
Implement error handling guidance in descriptions. Example: 'If the API returns rate_limit_exceeded, the tool will automatically retry with exponential backoff (max 3 attempts). If all retries fail, returns error_code=rate_limited with retry_after_seconds. If the user exceeds token budget, returns error_code=token_budget_exceeded and suggests splitting the input or increasing budget.'
Rename generic tools to follow verb-noun pattern. 'analyze' → 'analyze_code' or 'analyze_content'; 'chat' → 'chat_with_ai' or 'brainstorm'; 'challenge' → 'solve_challenge' or 'evaluate_solution'. This reduces ambiguity and aligns with the 90% baseline of A+ tools.
Add tool annotations (readOnlyHint, idempotentHint) where applicable. Per spec: all 16 tools are marked Risk=READ_ONLY, so ensure tool definitions include readOnlyHint: true in schema. This signals to agents that these tools can be retried safely without side effects.
Create a tools/ mapping document or update README to explain tool selection logic. E.g., 'Choose thinkdeep for single-model reasoning, consensus for multi-model synthesis, codereview for structural code analysis, debug for root-cause analysis, secaudit for security concerns. Choose one per request to avoid redundancy.'
Validate and expose input constraints for all parameters. If any tool accepts text, code, or file content, specify max length and character restrictions. Example: code_content: str = Field(description='Source code to analyze. Max 50,000 chars; larger files will be truncated.'); language: str = Field(description='Programming language (python|javascript|java|go|rust|cpp). Use to enable syntax-aware analysis.')