API server for converting academic paper content into professional visual schemas and rendered diagrams using AI models. Supports text input, PDFs, and images with multimodal AI processing.
This FastAPI server exposes 2 tools for academic diagram generation with serious definition quality gaps. Tool names lack clarity on what actually happens ('generate_schema' and 'render_image' are vague verbs). Descriptions are present but generic and lack LLM-optimized guidance on when/why to use each tool. Input schemas ARE present and typed (good), but parameter descriptions are superficial. No output schema documentation exists. Error handling is minimal, most errors will be HTTP exceptions without recovery guidance. The server lacks tool annotations (readOnlyHint, idempotentHint) despite both tools being read-only. No pagination, no result limits, and no indication of what the schema/image outputs actually contain structurally. Overall, this is a domain-specific tool wrapper that prioritizes API pass-through over agent-friendly design.
Step 1: The Architect. Generates a Visual Schema from paper content using a logic model. Supports text input, PDF pages, or images.
Step 3: The Renderer. Renders an academic diagram from the Visual Schema using OpenAI-compatible API with optional reference images for style guidance.
Tool names are generic action verbs with no object context. 'generate_schema' and 'render_image' do not clearly distinguish between an academic diagram schema and a rendered diagram. LLMs will struggle to infer when to call each. Better: 'generate_visual_schema_from_paper' and 'render_academic_diagram_from_schema'.
Parameter descriptions are minimal and lack LLM-actionable guidance. 'The visual schema with BEGIN/END PROMPT tags' is cryptic, what format? What do those tags mean? No example structure provided. 'Optional base64 encoded PDF pages or images' does not specify supported formats, size limits, or how many images are acceptable.
No output schema documentation. The tool descriptions say what goes in but never define what comes out. Does 'generate_schema' return a string? A structured object? Does 'render_image' return a base64 image, a URL, or an object with metadata? LLMs cannot plan downstream calls without knowing the response structure.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 47 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 29 | - | v1 |
No error handling guidance or recovery paths. If an API key is invalid, the server returns HTTPException(400) with detail='API Key is required.' But what if the API key is valid but the model name does not exist? Or the baseUrl is unreachable? The LLM gets a raw HTTP error with no indication of what to try next (retry? Use a different model? Check credentials?).
No tool annotations. Both tools are read-only (they generate outputs from inputs without modifying state), but the MCP registration includes 'Risk: READ_ONLY' as a comment, not a machine-readable annotation. The MCP spec expects toolAnnotations like 'readOnlyHint=true' or 'idempotentHint=true' in the tool registration.
API key exposure risk. Both tools accept 'config.apiKey' as a parameter, which means the API key is passed through the MCP request and likely logged in traces. Per the security pattern, credentials should NEVER be tool parameters, they should be injected server-side via environment variables or a vault. An agent calling this tool will embed the API key in its reasoning, traces, and request logs.
Missing dependency documentation. The 'config' parameter includes 'baseUrl' and 'modelName', but the descriptions do not explain the relationship between these fields and the actual LLM calls. Is 'baseUrl' a Gemini endpoint? An OpenAI-compatible endpoint? The code checks 'is_gemini_endpoint(url)' internally, but the tool description does not document this conditional behavior or tell the LLM which URLs to use.
No result limits or pagination. The tools process arbitrary amounts of input (PDFs with many pages, large images) and generate large outputs without documented constraints. A 500-page PDF converted to images will explode memory and token usage. The tool descriptions should specify 'max 50 pages per PDF' or similar, and the implementation should enforce it.