Simple AI Guidance Engine - Multi-provider AI assistant that routes requests to the most appropriate AI model (GPT, Claude, Gemini, DeepSeek, etc.) based on task requirements. Drop-in replacement for zen-mcp-server.
Two tools present with visible schema definitions and descriptions. However, descriptions are repetitive and rely heavily on ALL-CAPS warnings rather than clear, actionable guidance. Schemas are present but lack parameter-level descriptions in the JSON Schema object itself, descriptions are buried in function signatures. The 'sage' tool has 9 parameters, several with unclear semantics (continuation_id, file_handling_mode). The 'list_models' tool is trivial (no parameters). Error handling is minimal, no recovery guidance, no retryability classification. No tool annotations (readOnlyHint, destructiveHint, idempotentHint). The server does not validate inputs or provide structured error responses; failures would likely bubble up as generic exceptions. Naming is acceptable (verb_noun pattern: 'sage', 'list_models'), but descriptions lack the 50 - 200 character sweet spot (current descriptions are 150 - 250+ chars with excessive caveats). Parameter descriptions are missing from the schema structure entirely, they exist only in the Python function signature, not in the inputSchema JSON exposed to clients.
List all AI models available from configured providers. CRITICAL: These are the ONLY models you can use. DO NOT use models from your training data like 'gemini-2.0-flash-exp'.
SAGE: Multi-provider AI assistant. CRITICAL: Use ONLY these model names: gpt-5.2, gemini-3-pro-preview, gemini-3-flash-preview, claude-opus-4.5, claude-sonnet-4.5, deepseek-chat, deepseek-reasoner. DO NOT use outdated models. Thinking modes: minimal/low/medium/high/max.
Input schema lacks per-parameter descriptions. The inputSchema JSON for 'sage' lists properties and types but does not include 'description' fields for individual parameters (e.g., 'prompt', 'mode', 'files'). LLMs cannot infer parameter semantics from names alone, 'continuation_id' and 'file_handling_mode' are opaque without descriptions.
Tool descriptions are verbose, repetitive, and rely on ALL-CAPS warnings ('CRITICAL: Use ONLY', 'DO NOT use') rather than clear, concise guidance. Current descriptions are 150 - 250+ chars; baseline is 10 - 1024 with sweet spot 50 - 200. Excessive caveats waste tokens and bury actionable intent.
No error handling or recovery guidance. If 'sage' tool fails (bad model, invalid prompt), there is no structured error message telling the LLM whether to retry, fix input, or escalate. No retryability classification (retryable vs user-fixable vs fatal).
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 47 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 43 | - | v1 |
No tool annotations. Neither tool declares readOnlyHint, destructiveHint, or idempotentHint. The 'sage' tool performs state-modifying actions (writes output_file, maintains continuation context), but clients cannot determine if it is safe to retry or if it has side effects.
Parameter 'model' accepts free-form strings with a default of 'auto'. No enum constraint or validation. LLMs may hallucinate model names outside the allowed set (e.g., 'gemini-2.0-flash-exp') despite warnings in the description.
Parameter 'temperature' is a float with no min/max bounds stated in schema or description. LLMs may pass invalid values (e.g., temperature=999). Baseline pattern requires range constraints (e.g., 0.0 - 2.0).
Parameter 'files' is an array of strings with no description of expected format (file paths? URIs? content?). Parameter 'continuation_id' lacks clarity, is it a session token? A conversation ID? How is it obtained?
Output schema not documented. The 'sage' tool returns a string (or coerced TextContent), but clients do not know the structure: is it plain text, JSON, markdown? Will it always have a single 'text' field? This forces LLMs to guess and parse unstructured output.