The server defines 6 tools with explicit schemas and descriptions. All tools have descriptions (10 - 195 chars), which meet minimum length requirements. Input schemas are present and properly structured using Pydantic BaseModel. However, several definition quality issues reduce the overall score: (1) tool naming combines concerns ('ceo_and_board' violates single-responsibility; should be 'send_prompt_to_board' + 'synthesize_board_responses'); (2) parameter descriptions are adequate but lack actionable constraints (no enums for provider prefixes, no regex for file paths, no min/max for model lists); (3) output schemas are not documented in tool definitions, responses are ad-hoc TextContent with no structured field declaration; (4) error handling is absent, no guidance for retryability, user-fixable vs fatal errors, or recovery steps; (5) two tools (prompt, list_models) are inferred from code snippets rather than explicitly visible in the main server.py registration block, making full validation impossible.
Send a prompt to multiple 'board member' models and have a 'CEO' model make a decision based on their responses. IMPORTANT: You MUST provide absolute paths (e.g., /path/to/file or C:\path\to\file) for both file and output directory, not relative paths.
List all available models for a specific LLM provider
List all available LLM providers
Send a prompt to multiple LLM models
Send a prompt from a file to multiple LLM models. IMPORTANT: You MUST provide an absolute file path (e.g., /path/to/file or C:\path\to\file), not a relative path.
Send a prompt from a file to multiple LLM models and save responses to files. IMPORTANT: You MUST provide absolute paths (e.g., /path/to/file or C:\path\to\file) for both file and output directory, not relative paths.
Tool 'ceo_and_board' violates single-responsibility principle. Name combines two distinct actions: sending prompts to board models AND synthesizing responses via CEO model. Should be split into 'send_prompt_to_models' + 'synthesize_model_responses' or similar.
Parameter 'models_prefixed_by_provider' lacks an enum constraint or pattern validation. Accepts free-form strings like 'openai:gpt-4o' or 'o:gpt-4o', but LLMs may hallucinate invalid provider/model combinations (e.g., 'aws:claude'). Should specify valid provider prefixes as enum and document model availability via list_providers/list_models.
File path parameters (abs_file_path, abs_output_dir) lack path validation constraints. Descriptions state 'must be absolute path', but no regex pattern or length limits are declared. LLMs cannot verify validity before calling. Should add pattern constraint like '^(/|[A-Z]:\\)' and validation error guidance.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 43 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 35 | - | v1 |
No output schema documentation. Tools return TextContent with free-text responses (e.g., 'Model: X\nResponse: Y'). LLMs cannot parse the structure, they must extract model names and responses via string parsing. Should return structured JSON with fields like {model: string, response: string, status: string}.
No error handling guidance. Tool implementations call external APIs (OpenAI, Google, Groq, Ollama, Anthropic) without documented error responses. LLMs have no recovery path for: API auth failures, rate limits, model not found, file not readable, invalid provider. Should return error objects with retryability classification and next steps.
Tool descriptions lack discovery hints. 'Send a prompt to multiple LLM models' does not explain when to use prompt vs prompt_from_file vs prompt_from_file_to_file. Users/agents cannot infer which tool applies to their scenario without reading implementations.
ListProvidersSchema accepts no parameters but description is vague ('List all available LLM providers'). Should clarify: does it return provider names only? With model counts? With capability metadata (supports streaming, vision, etc.)? Response structure undefined.
Parameter 'ceo_model' has a default (DEFAULT_CEO_MODEL) but that constant is imported and unknown to the caller. Should document the exact default value in the schema description, e.g., 'default: anthropic:claude-3-5-sonnet'.
Logging configured at INFO level in server.py, but per-request logLevel control (io.modelcontextprotocol/logLevel in _meta) not implemented. Modern MCP servers should allow clients to request DEBUG/TRACE logging per call without restart.