MCP server for Napkin AI visual generation API with multi-storage support
The server defines 5 tools with complete input schemas and reasonable descriptions. Tool naming follows verb_noun convention (generate_*, check_*, download_*), which is correct. However, several tools combine multiple responsibilities (generate_and_wait, generate_and_save), and descriptions lack specificity about error recovery and prerequisites. All parameters are typed and most are described, but descriptions are generic and lack actionable guidance for LLMs. Output schemas are documented at the tool registration level but lack depth regarding field semantics and error cases. The generate_visual parameter schema is comprehensive (15+ optional fields) but overly complex, suggesting poor decomposition.
Check the status of a visual generation request. Returns progress information and file details when completed.
Download a generated visual file as base64-encoded data. Use the URL from check_status response's generated_files array.
Generate a visual and save to configured storage. Combines generation with automatic storage based on configured provider (local, S3, Google Drive, Slack, etc.).
Generate a visual and wait for completion. Combines generate_visual and polling check_status into a single operation.
Submit a visual generation request to Napkin AI. Returns a request ID for tracking. Use check_status to poll for completion.
Multiple tools combine unrelated responsibilities: generate_and_wait and generate_and_save both include generation logic plus a secondary concern (polling or storage). This violates single-responsibility principle and forces LLMs to decide which variant to use.
Parameter schema reuse across generate_visual, generate_and_wait, and generate_and_save (all with identical 15+ parameter lists) suggests the CommonInputSchema should be decomposed. If an agent always needs to specify format, visual_id, number_of_visuals, etc., the interface is too complex. Consider accepting a struct or config object, or providing preset variations.
Descriptions do not explain error recovery or prerequisites. For example, check_status does not document what statuses are possible (pending, processing, completed, failed?), what to do if a request has expired, or how long URLs remain valid. LLMs cannot infer retry logic or fallback strategies.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 56 | 2025-06-18+ | v2 |
| 2026-03-09 | F | 37 | 1.0.0+ | v1 |
Output schemas are documented at registration but lack detail. The code shows outputSchema for generate_visual (id, status, warning) but does not show schemas for check_status or download_visual. Missing documentation of what fields indicate completion, error states, and what an LLM should do next if a request fails.
No tool provides dry_run output or validation feedback. The dry_run parameter exists (intended to 'Validate inputs without calling the API'), but the output schema and description do not explain what a dry_run response looks like or how LLMs should interpret it.
Tool descriptions do not disambiguate when to use generate_visual vs generate_and_wait vs generate_and_save. An LLM presented with three generation tools must guess which is appropriate. generate_visual is for async workflows? generate_and_wait for sync? generate_and_save only if storage is pre-configured? This is not stated.
Parameter constraints are not enforced or validated in descriptions. For example, number_of_visuals has min=1, max=4, but the description does not state what happens if an LLM passes 5 or 0. Similarly, width/height have min=100, max=10000, but no description of why or what error to expect.