MCP server that communicates with Ollama to provide task decomposition, result evaluation, and local LLM-powered text/JSON generation capabilities
This MCP server exhibits several critical quality gaps across naming, descriptions, parameter documentation, and schema clarity. While all 4 tools have basic schemas and descriptions, most fall short of production standards. Tool names lack consistent verb-noun patterns (decompose-task, evaluate-result are acceptable but ambiguous in intent). Parameter descriptions are minimal, many lack actionable constraints or context. No output schemas are documented. Error handling exists but is generic. The server appears to prioritize basic functionality over LLM-friendly tool design. Average per-tool score across the 4 tools is 42, placing this in the 'D' (Poor) range.
Add a new task
Break down a complex task into manageable subtasks
Evaluate a result against specified criteria
Execute an Ollama model with the specified parameters
Tool names lack consistency and clarity. 'decompose-task' and 'evaluate-result' use hyphens instead of underscores (non-standard for Python/JSON tooling) and are vague about what decomposition strategy is applied or what 'evaluation' means operationally. 'add-task' is acceptable but doesn't distinguish this tool from other task creation patterns.
Parameter descriptions are minimal and lack actionable constraints. Example: 'granularity' in decompose-task is described only as 'Granularity level for decomposition', it does not explain WHY the agent should choose 'high' vs 'medium' vs 'low', what output size each produces, or when to use each. Baseline for A+ tools is 100% of params have descriptions ≥50 chars with context.
No output schemas are documented. LLMs cannot plan downstream tool calls or extract required fields (e.g., subtask IDs for further decomposition) without knowing the response structure. All 4 tools lack documented return types.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 46 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 11 | - | v1 |
Tool 'run-model' has a required parameter 'prompt' but is missing 'model' (marked required=false, but the description says 'The model to execute', unclear whether model is optional or should default to a server-side configuration). Dependency between 'model', 'temperature', and 'max_tokens' is undocumented.
Parameter 'criteria' in evaluate-result is typed as object with additionalProperties (weights), but the description does not explain: (1) what criteria keys are valid, (2) what weight range is expected (0 - 1? 0 - 100?), (3) whether weights must sum to 1 or are independent. This forces LLMs to guess or fail.
Tool descriptions are generic and lack WHEN/WHY guidance. 'Add a new task' does not explain when to use this tool vs decompose-task, whether it supports immediate execution, what happens if a task with the same name already exists, or what task_id is returned for downstream use.
Error handling in source code uses a custom MCPServiceError class with generic JSON serialization, but error messages do not guide recovery. An LLM receiving 'Task not found' has no actionable next step, should it search for the task by name? List all tasks? Error responses must include recovery hints.
Numeric parameters lack bounds. 'max_tokens' in run-model has no min/max specified; LLMs could pass 999999 or negative values. 'priority' in add-task is a number with no range; is it 1 - 5? 0 - 100? Open-ended numbers invite hallucinated values.