Exposes PaperBanana's core functionality (academic diagram and plot generation) as MCP tools usable from Claude Code, Cursor, or any MCP client.
PaperBanana provides 10 well-named, verb-first tools with comprehensive parameter schemas and descriptions. However, critical gaps in output schema documentation, error handling guidance, and security considerations prevent a higher score. Tool names are clear (generate_diagram, batch_plots, etc.), and parameter descriptions are detailed (10-50 characters for each). Input schemas are fully visible with proper JSON Schema types. However, output schemas are not documented in the source code, responses are inferred from the implementation rather than formally declared. Error recovery guidance is absent: tools provide no hints about retryable vs. fatal errors, no suggestions for next steps on failure, and no validation of parameter ranges. Security is a concern: batch operation file paths accept raw user input without path traversal validation visible in the schema. The per-request _meta logLevel mechanism (current spec) is not evident in tool definitions. The 'Risk' annotations (WRITE, READ_ONLY) are helpful but toolAnnotations (readOnlyHint, destructiveHint) are not formally present in the MCP schema. Average tool definition quality is solid (clear naming, structured inputs), but production-grade error handling and output documentation are missing.
Batch generate methodology diagrams from a YAML/JSON manifest.
Batch generate statistical plots from a YAML/JSON manifest.
Continue a prior methodology diagram run with more refinement or user feedback.
Continue a prior statistical plot run with more refinement or user feedback.
Download the PaperBananaBench reference set (~298 examples) for use in local evaluation or as context.
Evaluate a generated diagram against a reference image using a VLM judge.
Output schemas not formally documented in tool definitions. Responses are inferred from implementation, not declared as structured types. LLMs cannot plan downstream tool calls without knowing what fields to expect.
No error handling guidance or recovery instructions. Tools do not indicate which errors are retryable, which require user action, or what the LLM should do next on failure. Error messages are not visible in the schema.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | B | 72 | <=2025-11-25 | v2 |
| 2026-03-09 | C | 68 | - | v1 |
Evaluate a generated plot against a reference image using a VLM judge.
Generate a publication-quality methodology diagram from text.
Generate a statistical plot (bar, line, scatter, etc.) from JSON data.
Generate a full-paper figure package (planning + optional generation) from a methodology section.
File path parameters (manifest_path, output_dir) lack validation constraints and path traversal protection visible in schema. No pattern, length limit, or allowed directory constraints documented. LLMs could pass arbitrary paths including '../' sequences.
Tool annotations (readOnlyHint, destructiveHint, idempotentHint) are not present in MCP schema, though Risk metadata is provided in comments. Tools lack formal semantic markers for LLM planning.
Numeric parameters (iterations, aspect_ratio, max iterations) lack explicit min/max bounds. LLMs could pass absurd values (iterations=10000, max auto_refine loops=unlimited). No constraint validation visible in schema.
batch_diagrams and batch_plots accept manifest file paths as free-form strings. No documented format, no example, no instruction for LLMs on how to construct these manifests or what fields are required.
No pagination or result limiting guidance for large batch operations. batch_diagrams and batch_plots offer parallel execution but no cap on concurrent jobs, result count, or guidance on splitting large manifests.