MCP server that gives online LLMs local hands: OpenAI-compatible local inference services (KoboldCpp, Unsloth, llama.cpp server, LM Studio, Ollama, ...) for cheap text labor and vision work (analyze / OCR / compare).
Three tools with comprehensive descriptions and well-structured schemas. local_run and local_vision have detailed, LLM-optimized descriptions (200+ chars) explaining WHEN to use them and what they do. All parameters have types and descriptions. However, output schemas are not explicitly documented in the code, only input schemas are visible. Error handling is present but recovery guidance is minimal. Tool names are clear and action-oriented (local_run, local_vision, local_status). No critical security issues detected (no secrets in params). Composition is sound: three focused tools with clear responsibilities.
Run ONE prompt on a LOCAL OpenAI-compatible text model on this machine (KoboldCpp / Unsloth / llama.cpp / LM Studio / Ollama ...). Use it for simple, repetitive, token-cheap labor instead of spending main-model tokens: batch rewrites, name translations, string munging, deduplication, short-text summarization, structured extraction, and other mechanical text work. The prompt is sent to the local model as a user message (the server applies its own chat template).
List configured backends and their health status.
Send one or more images to a LOCAL multimodal model (KoboldCpp / Unsloth / llama.cpp with an mmproj projector, ...) for image understanding: OCR / text extraction, describing or analyzing images, reading charts and screenshots, comparing images. Use it to offload vision work from the main model — especially when the main model is text-only and cannot see images itself: pass the image path (or a data:/http(s) URL) in image_paths / image_urls and the local vision model reads it. Images come from local file paths (image_paths), data:/http(s) image URLs (image_urls), or both. png/jpg/jpeg/webp/gif/bmp, up to 20 MB each. The local server must run a multimodal model with its mmproj projector loaded. Pick a mode for structured output, or pass your own prompt for a specific question (never both): - `analyze` (default): an '# Image Analysis Report' with 8 fixed sections: Summary / Image Metadata / Layout & Composition / Visible Text (VERBATIM) / Objects & Elements / People & Actions / Semantic Context & Inferences / Uncertainties & Gaps. - `ocr`: character-exact text extraction in reading order (the model does NOT 'fix' typos or drop symbols; unresolvable glyphs are noted). - `compare`: with 2-4 images in ONE call, an '# Image Comparison Report': Per-Image Summaries / Common Elements / Key Differences / Text Differences (VERBATIM) / Overall Conclusion. For independent per-image analysis use one call per image; use `compare` only when the task needs joint reasoning across images. The local model may also return reasoning text (thinking) before its answer; it is reported separately as `reasoning`. When relaying the local model's output, keep full fidelity: never rephrase, shorten, 'fix', or invent visual details the report did not return; preserve any uncertainty the report explicitly states.
Output schemas not documented. Code shows input schemas but no explicit return type definitions for any tool. LLMs cannot plan downstream calls or extract fields without knowing what local_run, local_vision, and local_status return.
local_status description is minimal ('List configured backends and their health status.'). Does not explain WHEN to call it, what fields are returned, or how to interpret health status. Baseline for param descriptions is 72 chars; this is 50.
Error handling present but lacks recovery guidance. Code raises LocalCallError and LocalError exceptions, but descriptions do not tell LLMs what to do on failure (retry? call local_status? ask user?). Pattern: recovery-guide.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 69 | 2026-07-28+ | v2 |
No tool annotations (readOnlyHint, destructiveHint, idempotentHint). All three tools are read-only and idempotent, but this is not declared in the schema. Agents cannot infer safety properties without explicit hints.
local_vision 'mode' parameter has a default ('analyze') but no enum constraint in the visible schema. Description lists valid modes (analyze, ocr, compare) but schema should enforce this formally to prevent LLM hallucination.