A FastMCP server that extracts LaTeX formulas from images using a vision language model
Single tool with a descriptive name but significant gaps in schema completeness, parameter documentation, and error handling. The tool definition is visible in mcp_server.py and explicitly registered via @mcp.tool(), so it is not inferred. Tool name 'extract_latex' is action-verb-prefixed and clear. Description is adequate (97 chars, within 10-1024 range) but lacks guidance on when to use it, prerequisites, or what to expect. Input schema includes two parameters (image_base64, prompt) but both lack formal type constraints and the prompt parameter's default=null introduces ambiguity. No output schema is documented, LLMs cannot reason about what structure to expect. No error handling guidance is provided, failures would require the LLM to guess remediation steps. Security is reasonable (no secrets in params, read-only operation) but lacks explicit scope declaration. Overall, this is a functional but underdeveloped tool suitable for a single-purpose prototype, not production use.
Extracts LaTeX mathematical formulas from images using a vision language model
Output schema not documented. LLMs cannot infer what the tool returns. Code shows it returns extracted LaTeX as a string or None, but this is not declared in the tool signature or description.
Parameter descriptions are minimal and lack format/constraint guidance. 'image_base64' describes what it is but not required encoding or max size. 'prompt' is optional but defaulting to null leaves unclear what happens when omitted, does the tool use a hardcoded fallback prompt?
No error handling guidance. Tool can fail silently (extract_answer returns None if no regex matches) or raise exceptions (PIL, OpenAI API calls). LLM has no recovery path, cannot know whether to retry, adjust input, or inform user.
Tool description lacks prerequisite or when-to-use guidance. It mentions 'vision language model' but does not state that an OpenRouter API key and internet connectivity are required, or that the model is 'qwen/qwen2.5-vl-72b-instruct:free'.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 46 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 40 | - | v1 |
prompt parameter defaults to None with no documented fallback. Code shows if prompt is None, an empty string is concatenated, resulting in a generic prompt. This should be explicit in the schema and description.
No confirmation or dry-run capability. Tool sends images to external VLM and may log results. For sensitive images (exam papers, confidential documents), there is no confirmation step, destructive/data-leakage risk is not acknowledged.