MCP server that exposes local Ollama instances as tools for Claude Code
The server exposes two well-named tools (ollama_generate, ollama_chat) with clear verb_noun naming and contextual descriptions. However, the descriptions are relatively generic and lack some critical guidance for agent selection. Parameter schemas are mostly complete with types and defaults, but lack formal constraints (enums, min/max bounds) that would prevent invalid input. Error handling is minimal, the tools return plaintext error strings rather than structured, actionable recovery guidance. No tool annotations (readOnlyHint, etc.) are declared. Output format is documented in prose but not as a formal schema. The implementation shows good engineering practices (config loading, retry logic, version checking) but falls short of production-grade tool definition standards.
Send a multi-turn conversation to a local Ollama model. Use this when you need to have a back-and-forth with the local model, or when prior context matters for the response.
Send a prompt to a local Ollama model and return the response. Use this for code generation, documentation drafts, quick questions, and tasks that don't require frontier-model reasoning.
Parameter constraints missing for host and model. 'host' accepts 'local' or 'server' but is not declared as an enum; 'model' accepts any string but defaults to empty, inviting hallucinated model names. LLMs will guess invalid values.
Output schema not formally documented. Tool descriptions mention 'response' and 'metadata' fields (tokens, tok/s) in prose, but return type is declared as 'str', a plaintext string. LLMs cannot parse field structure from a string return type and cannot extract data for downstream tool calls.
Error handling returns plaintext strings ('Error: host must be one of...') instead of structured, actionable recovery guidance. Errors lack categorization (retryable vs user-fixable vs fatal) that agents rely on to decide whether to retry, ask the user, or escalate.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 67 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 43 | - | v1 |
Tool descriptions do not explain WHEN to use each tool instead of the other. Both ollama_generate and ollama_chat accept similar parameters and both query Ollama. The description for ollama_generate says 'for code generation, documentation drafts, quick questions' but ollama_chat could do the same. Ambiguous differentiation forces LLM to guess.
No tool annotations declared. Tools do not declare readOnlyHint (both are read-only), idempotentHint (both are idempotent), or destructiveHint (neither is destructive). Agents cannot distinguish safety properties.
Numeric parameters (timeout) lack bounds. timeout is a float with no min/max specified. LLMs could pass negative, zero, or absurdly large values (1e9 seconds) without validation.
Parameter descriptions lack format and constraint hints. The 'messages' parameter in ollama_chat is described as 'List of message dicts with role and content keys' but does not specify: what are valid values for 'role'? (user, assistant, system?). Are keys case-sensitive? What happens if required keys are missing? LLMs must guess or trial-and-error.