MCP server for delegating mechanical tasks to local LLMs via Ollama
The server defines 7 tools with clear, well-structured schemas and good delegation guidance embedded in descriptions. Tool naming follows verb_noun convention (local_summarize, local_draft, local_classify, local_extract, local_transform, local_complete, local_status). All tools have input schemas with proper JSON Schema type definitions and parameter descriptions. However, there are notable gaps: (1) NO documented output schemas, tools return unstructured responses from an LLM, with no clear guidance on what structure the agent should expect; (2) Parameter descriptions lack explicit constraints (e.g., max_length in local_summarize says 'Approximate max words' but doesn't specify if there's a hard limit or sensible bounds); (3) Error handling is absent, no guidance on recovery if the local LLM service is down or returns malformed JSON; (4) No per-tool output documentation means agents must guess at response structure. Strengths: tool names are distinct and action-oriented; descriptions include explicit delegation guidance stating when to use each tool vs. alternatives (e.g., 'Use this when you need to summarize content that doesn't require your full reasoning'); all required parameters are marked; optional parameters are well-documented (style, max_length, focus in summarize).
Classify text into categories using a local LLM. DELEGATION GUIDANCE: Use for sorting, tagging, organizing content. Good for batch classification tasks where the categories are clear-cut.
Raw completion using local LLM for maximum flexibility. DELEGATION GUIDANCE: Use when other tools don't fit. You control the prompt entirely. Good for custom tasks that don't match predefined patterns.
Generate an initial draft using a local LLM that you can then refine. DELEGATION GUIDANCE: Use this for boilerplate content, initial drafts, template-based generation. You (Claude) should review and refine the output - the local model does the grunt work, you do the quality control.
Extract structured information from text using a local LLM. DELEGATION GUIDANCE: Use for parsing documents, extracting specific fields, converting unstructured text to structured data.
Check the status of the local LLM and available models.
No documented output schemas. Tools return raw LLM responses (strings or parsed JSON), but there is no specification of what fields or structure the agent should expect. This forces agents to parse unstructured text or assume response format, increasing error likelihood.
Parameter constraints lack explicit bounds. max_length in local_summarize is described as 'Approximate max words (default: 150)' but does not specify minimum, maximum, or hard limits. temperature in local_complete lacks explicit range (should be 0.0-2.0). agents may pass invalid values.
No error handling guidance. If the local LLM service is unreachable, returns malformed JSON, or times out, there is no documented recovery path or error classification (retryable vs. fatal). Agents will receive raw errors with no actionable next steps.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 61 | 2026-07-28+ | v2 |
| 2026-03-09 | C | 60 | - | v1 |
Summarize text using a local LLM. Use for long documents, research notes, or any content that needs condensing. DELEGATION GUIDANCE: Use this when you need to summarize content that doesn't require your full reasoning - bulk file summarization, extracting key points from research, condensing meeting notes.
Transform text according to instructions using a local LLM. DELEGATION GUIDANCE: Use for formatting changes, style conversions, simple rewrites. Good for mechanical transformations that don't require deep reasoning.
local_status description is vague ('Check the status of the local LLM and available models') and lacks guidance on what fields it returns or when to call it. Input schema is empty (no parameters), which is correct, but output structure is not documented.
No batch variant for multi-item operations. If an agent needs to classify 10 texts, it must call local_classify 10 times sequentially. A batch mode (classify_many) would be more efficient and reduce token overhead.