MCP server for orchestrating Ollama models with multi-step workflows, deep reasoning, planning, consensus building, code review, and debugging capabilities
The server has 9 tools with mostly complete schemas and reasonable descriptions, but exhibits several quality gaps: (1) Tool names lack action-verb prefixes (chat, debug, thinkdeep, planner are vague; should be something like 'start_chat', 'run_debug', 'analyze_deeply', 'create_plan'). (2) Descriptions vary in quality, some are good (chat: 144 chars, codereview: 161 chars), but others are verbose and repetitive (consensus: 245 chars, planner: 240 chars, exceeding the 200-char baseline for dense information). (3) Most parameters have descriptions and types, but some optional parameters lack clarity on defaults and interdependencies (e.g., consensus.models requires minimum 2 items but the relationship between current_model_index and which stance to use is undocumented). (4) Output schemas are NOT documented in the tool definitions, responses are inferred from Ollama API behavior rather than explicitly declared, violating the 'document output schema' requirement. (5) Error handling is generic ('Error executing tool X: {message}') without recovery guidance or classification. (6) The step-based workflow tools (debug, thinkdeep, planner, consensus, codereview, precommit) have similar parameter shapes and descriptions, creating cognitive overhead for LLMs to distinguish when to use each.
General chat and collaborative thinking partner for brainstorming, development discussion, getting second opinions, and exploring ideas. Use for ideas, validations, questions, and thoughtful explanations.
Performs systematic, step-by-step code review with expert validation. Use for comprehensive analysis covering quality, security, performance, and architecture. Guides through structured investigation to ensure thoroughness.
Builds multi-model consensus through systematic analysis and structured debate. Use for complex decisions, architectural choices, feature proposals, and technology evaluations. Consults multiple models with different stances to synthesize comprehensive recommendations.
Performs systematic debugging and root cause analysis for any type of issue. Use for complex bugs, mysterious errors, performance issues, race conditions, memory leaks, and integration problems. Guides through structured investigation with hypothesis testing and expert analysis.
Shows available Ollama models, their names, and capabilities.
Tool names lack action-verb prefixes. 'chat', 'debug', 'thinkdeep', 'planner', 'consensus' are nouns/adjectives; LLMs infer intent from verb-noun patterns. Should be 'start_chat' or 'initiate_conversation', 'run_debug' or 'diagnose_issue', 'analyze_deeply', 'create_plan', 'build_consensus'.
Output schemas are NOT documented. The tool definitions declare only input schemas (via Zod). Responses are unstructured, LLMs cannot predict the result structure to plan downstream calls or extract relevant data. Pattern requires 'Document the output schema.'
Error handling is generic without recovery guidance. All tools catch errors and return 'Error executing tool X: {message}'. Does not classify errors as retryable/user-fixable/fatal, nor suggest next steps. Pattern requires 'Error responses must tell the LLM what to do next.'
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 49 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Breaks down complex tasks through interactive, sequential planning with revision and branching capabilities. Use for complex project planning, system design, migration strategies, and architectural decisions. Builds plans incrementally with deep reflection for complex scenarios.
Validates git changes and repository state before committing with systematic analysis. Use for multi-repository validation, security review, change impact assessment, and completeness verification. Guides through structured investigation with expert analysis.
Performs multi-stage investigation and reasoning for complex problem analysis. Use for architecture decisions, complex bugs, performance challenges, and security analysis. Provides systematic hypothesis testing, evidence-based investigation, and expert validation.
Get server version, configuration details, and list of available tools.
Multi-step workflow tools (debug, thinkdeep, planner, consensus, codereview, precommit) have near-identical parameter signatures (step, step_number, total_steps, next_step_required, findings, model, continuation_id, etc.). This creates confusion for LLMs, no clear distinction in interface. Each should have domain-specific parameters or clear UX differences.
Descriptions exceed 200-char baseline for three tools (consensus: 245, planner: 240, debug: ~220). Verbose descriptions waste tokens and bury key details. Should distill to 'WHAT, WHEN to use it, WHEN to use it instead of similar tool' in 100-150 chars.
Optional parameters lack interdependency documentation. E.g., consensus.models requires minimum 2 items, but if current_model_index=0, which model's stance applies? How does stance_prompt override the default stance? Undocumented dependencies cause silent misuse.
listmodels and version tools have empty input schemas ({}). While correct, their descriptions lack clarity on what structure is returned. 'Shows available Ollama models, their names, and capabilities' does not indicate pagination, field names, or how to extract a specific model. Should document output shape.
chat tool accepts optional 'files' and 'images' parameters as arrays of strings, but descriptions are generic. No guidance on supported formats, paths vs URLs, max file size, or what happens if a file doesn't exist. Pattern requires actionable format specs.