An MCP server for translating PDF documents with mathematical content using various translation services
The server exposes only 1 tool (translate_pdf) with a basic description and partial schema. The tool description is present (114 chars, within baseline 10-1024 range) and describes the action clearly. However, critical gaps emerge: (1) Input parameters lack descriptions in the JSON schema, the schema object shows only type and description fields for each param, but these are minimal ('absolute path of input pdf', 'language to translate from', 'language to translate to') and do not specify constraints, valid language code formats, or error recovery guidance. (2) No output schema is documented, the function returns a plain string with file paths, but the agent cannot parse structured results or extract metadata. (3) No error handling or recovery guidance, the tool silently fails if the file doesn't exist, language codes are invalid, or translation fails. (4) Parameter validation is absent from the visible code, no enums for lang_in/lang_out despite 'google translate lang_code' being mentioned. (5) Tool naming is acceptable (verb_noun: 'translate_pdf') but the description conflates multiple concerns (reading file, translating, writing outputs) without splitting into composable steps. (6) The description mentions 'google' service hardcoded but offers no guidance on supported languages or error cases. Overall, this is a minimal tool definition that would not pass production code review.
translate given pdf. Argument `file` is absolute path of input pdf, `lang_in` and `lang_out` is translate from and to language, and should be like google translate lang_code. `lang_in` can be `auto` if you can't determine input language.
Input schema parameters lack detailed descriptions and constraints. The schema includes basic type and description, but does not specify enum values for lang_in/lang_out, valid language code formats (e.g., ISO 639-1), minimum/maximum lengths, or error recovery steps. Parameter descriptions should explain what language codes are accepted (e.g., 'ISO 639-1 language code like "en", "zh", "fr"; pass "auto" to auto-detect source language').
No output schema documented. The tool returns a plain string with file paths concatenated as plaintext. Agents cannot structure-parse the result or extract the mono/dual PDF paths programmatically. Return a JSON object with fields like {mono_path: string, dual_path: string, status: string} to enable downstream tool composition and proper error handling.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 45 | <=2025-11-25 | v2 |
| 2026-03-09 | D | 52 | - | v1 |
No error handling or recovery guidance. The tool will fail silently if: (1) file does not exist, (2) lang_in/lang_out are invalid language codes, (3) translation service is down, (4) output directory is not writable. Each error case should return a distinct, actionable message (e.g., 'File not found: /path/to/file. Ensure the path is absolute and the file exists.' or 'Invalid language code: "xyz". Supported codes: en, zh, fr, ... (use "auto" to detect source language).') to guide the LLM's next step.
Tool description lacks key context. The description says 'google translate lang_code' but does not explain: (1) which language codes are supported, (2) whether 'auto' detection works for all inputs, (3) expected output file naming convention (mono vs dual), (4) what to do if translation fails. Add examples: 'Translates PDF text from one language to another using Google Translate. Supported lang_in and lang_out codes: en, zh, fr, de, ja, ko, ... (full list at https://...); lang_in can be "auto" to auto-detect. Returns two PDFs: one with only translated text (mono) and one with original + translation (dual). If translation fails due to invalid codes or service unavailability, returns an error message.'
Missing parameter enum constraints. lang_in and lang_out should be enums of supported language codes rather than free-form strings. The description mentions 'google translate lang_code' but does not provide the enum list in the schema. This invites the LLM to hallucinate unsupported codes like 'english' or 'chinese' instead of 'en' or 'zh'. Add enum: ["en", "zh", "fr", "de", "ja", "ko", ...] for lang_out; add ["auto", "en", "zh", ...] for lang_in to reflect the 'auto' option.
Tool combines multiple concerns without composability. The translate_pdf function reads the file, translates, writes two output PDFs, and logs progress, all in one step. This prevents agents from controlling individual steps (e.g., previewing translation before writing, choosing which output format to save). Consider splitting into: (1) translate_pdf_stream(file, lang_in, lang_out) → returns structured {mono_content, dual_content, metadata}, (2) save_translated_pdf(content, output_path, format) → writes file. This enables agents to compose steps and handle failures granularly.
File parameter accepts absolute paths without validation. The description says 'absolute path of input pdf' but does not validate: (1) that the path is absolute (not relative), (2) that the file exists, (3) that the file is a valid PDF, (4) that the file is readable by the server. Add validation: 'file must be an absolute path (starting with / on Unix or C:\ on Windows), must exist and be readable, and must be a valid PDF. Returns an error if any constraint is violated.'