MCP server for OCR, text extraction, language detection, translation, and email sending on images using FastMCP
This server has significant definition quality gaps. While tool names follow basic verb_noun conventions, most parameter descriptions are minimal or missing entirely. Input schemas are present but sparse, many parameters lack detailed constraint documentation. Output schemas are either undocumented or return free-form strings/dicts without typed field definitions. Error messages are generic ('ERROR: ...') rather than actionable recovery guidance. The toolset also includes several single-purpose utilities (strings_to_chars_to_int, create_thumbnail) that feel tangential to the core image analysis domain. No evidence of parameter enums, validation hints, or format constraints in descriptions. The send_email tool exposes environment variable names in error messages rather than sanitizing them. Overall, this reads like an educational/prototype server rather than production-grade tooling.
Create a thumbnail from an image
Detect language of the given text and return ISO language code.
Perform OCR on a local image path and return extracted text as a clean string. Removes newlines and unnecessary whitespace but preserves accented characters.
Perform OCR and return {'text': str, 'avg_confidence': float}. Removes newlines and unnecessary whitespace but preserves accented characters.
Send an email using Gmail SMTP.
Return the ASCII values of the characters in a word
Lightweight summarizer: truncate to max_length or return first sentence(s).
Minimal parameter descriptions across all tools. Example: ocr_image 'languages' parameter is described as 'OCR languages in Tesseract format (default: 'eng+spa')', no guidance on format, valid syntax, or when to use alternatives. detect_language has no parameter description at all. This violates the 'every parameter needs a description' rule and forces LLMs to guess input formats.
No output schemas documented. Tools return free-form strings or dicts without specifying field types or structure. Example: ocr_with_confidence returns dict with 'text' and 'avg_confidence' fields, but return type annotation shows only 'dict', LLMs cannot plan downstream operations without knowing the response structure. translate_text returns a bare string with no metadata about source language, confidence, or character count that might be useful for chaining.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 46 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 40 | - | v1 |
Translate text to destination language using deep_translator.
Error handling is non-actionable. Tools return strings like 'ERROR: pytesseract not available on server.' or 'ERROR: Failed language detection: {e}', no guidance on recovery steps, whether to retry, or what to try next. Per pattern:recovery-guide, errors should tell the LLM what to do: 'OCR failed: Tesseract binary not found. Install pytesseract and ensure Tesseract is in PATH, or call list_supported_languages() to verify service health.'
Security: send_email tool exposes environment variable names (GMAIL_ADDRESS, GMAIL_APP_PASSWORD) in error responses. Response like {'status': 'error', 'error': 'GMAIL_ADDRESS not set in environment'} leaks configuration details to logs and context. Per pattern:secret-injection, error messages should never reveal secret names or configuration structure, say 'Authentication failed. Contact administrator.' instead.
Undocumented parameter constraints and enums. Example: translate_text 'dest' parameter accepts language codes but no enum is declared and description offers no list of valid values ('en', 'es', 'fr', etc.). LLMs will hallucinate invalid language codes. ocr_image 'languages' parameter mentions 'Tesseract format' but does not specify syntax (is it 'eng+spa' or 'eng spa' or 'en,es'?). Per pattern:constrained-input, free-form strings should be replaced with enums or formal constraints.
Tool composition includes tangential utilities misaligned with domain focus. strings_to_chars_to_int and create_thumbnail are disconnected from image text analysis. These belong in a separate utilities server, not bundled here. Per pattern:tool, each tool should do exactly one thing within a coherent domain. Bundling unrelated tools dilutes the server's purpose and forces LLMs to reason about irrelevant options.
No input validation hints in descriptions. Example: summarize_text 'max_length' has no range guidance (minimum 1? maximum 10000?). send_email 'to', 'subject', 'body' have no length limits or format hints (email regex, subject max length, body encoding). Per review:param-validation-rules, descriptions should specify ranges and expected formats so LLMs can validate before calling.
detect_language and translate_text lack parameter descriptions entirely. detect_language 'text' parameter has no description. translate_text 'text' parameter has no description (only 'dest' does). This violates the 'every parameter needs a description' rule and leaves LLMs guessing whether these fields accept strings, arrays, or structured objects.
Tool descriptions are too brief and lack context for LLM selection. Example: create_thumbnail: 'Create a thumbnail from an image', does not answer WHEN to call it vs ocr_image, what size it produces, what format it returns, or whether it's a step toward downstream analysis. summarize_text: 'Lightweight summarizer: truncate to max_length or return first sentence(s).', vague about behavior (does it return first N sentences or exactly 2?). Per pattern:tool-description, descriptions should state WHAT, WHEN, and prerequisites.