Local GPU translation MCP server — TranslateGemma via Ollama, 57 languages, zero cloud dependency
Polyglot MCP demonstrates solid definition quality with comprehensive parameter schemas and clear descriptions for all 6 tools. All tools use Zod schema validation with proper type definitions and descriptions. Tool names follow verb_noun convention (translate, list_languages, etc.). Descriptions are specific and contextual (100-250 chars average), explaining what each tool does and when to use it. However, there are gaps in output schema documentation, error handling guidance, and parameter constraint clarity that prevent a higher score.
Verify Ollama is running and TranslateGemma is installed.
List all 57 languages supported by TranslateGemma for translation.
Translate text between any of 57 supported languages using TranslateGemma running locally on your GPU via Ollama. Automatically starts Ollama and pulls the model if needed. Includes a built-in software glossary for accurate technical translations.
Translate markdown content into multiple languages at once (default: 7 languages — Japanese, Chinese, Spanish, French, Hindi, Italian, Portuguese). Runs translations concurrently with GPU-safe semaphore limiting. Returns translated content for each language with optional nav bar injection.
Translate a markdown document while preserving its structure. Code blocks, HTML elements, URLs, badges, and table formatting are kept intact. Only prose content (headings, paragraphs, taglines, table cells) is translated. Supports segment-level caching.
Output schemas not documented. Tool descriptions do not specify what fields are returned or their types. For example, translate() does not document whether it returns {text, language_pair, model_used, cache_hit} or just the translated text string. LLMs cannot plan downstream operations without knowing response structure.
Error handling lacks recovery guidance. Tool descriptions mention Ollama dependency and model pulling but do not explain what happens on failure (e.g., 'If Ollama is not running, the tool will attempt to start it. If startup fails, call check_status first to diagnose.') or how to recover from specific failures (network, GPU memory, invalid language code).
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 68 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 53 | - | v1 |
Translate a README.md FILE into all 7 supported languages and WRITE the README.<lang>.md files to disk (ja, zh, es, fr, hi, it, pt-BR), refreshing the language nav bar in the source and each translation. Runs entirely locally on the GPU via Ollama. Returns a STATUS SUMMARY — per-language result.
Parameter constraints insufficiently documented. The 'model' parameter description lists options ('translategemma:27b', ':12b', ':4b') but does not explain the trade-offs clearly (quality vs speed vs VRAM requirements). The 'concurrency' parameter allows 'default: 2, max: 3' but does not explain why or what happens if user provides 4+. Parameter descriptions should state format, range, and enforcement.
Language parameter accepts both codes and names ('en' or 'English') but the interaction is not fully specified. If an LLM passes an invalid code like 'engish' or unsupported language like 'Klingon', the error message and recovery path are not documented. Tool should validate early and return actionable errors like 'Invalid language: got "engish". Did you mean: en (English)? Call list_languages() to see all 57 supported options.'
translate_readme writes files to disk but does not document the exact file paths, naming scheme (README.<lang>.md), or what happens if files already exist (overwrite? error? backup?). The description says 'WRITE the README.<lang>.md files' but does not specify: working directory context, permissions required, or rollback behavior on partial failure.
No documentation of pagination or result limits. translate_markdown and translate_all do not specify max markdown size, max output length, or token limits. If an LLM passes a 100KB markdown file, does the tool accept it? Return it in full? Truncate? Stream? Timeout? Without limits documented, agents cannot estimate feasibility and may hit silent failures.