MCP server for converting articles to Xiaohongshu (小红书) style image cards with AI cover generation
The server defines 4 tools with explicit schemas and descriptions. All tools have type-constrained parameters (enums for theme, ratio, fontSize). However, there are notable gaps: (1) output schemas are completely undocumented, no indication of what text_to_images or file_to_images actually return (image URLs? file paths? base64?); (2) parameter descriptions exist but are minimal (8-20 chars on average, well below the 72-char baseline); (3) no error handling guidance, what happens if Gemini fails? If file doesn't exist? If rendering times out?; (4) tool names are clear and verb-forward (text_to_images, file_to_images, estimate_pages, list_themes) but composition could be tighter, estimate_pages and list_themes are utility tools that could be combined; (5) no idempotency or side-effect warnings despite text_to_images having a WRITE risk flag and file_to_images supporting file operations.
Estimate how many pages the text will be split into without generating images
Convert a local markdown or text file to a series of images. Supports .md, .txt, .text files. Automatically extracts title from the first # heading in markdown files. Cleans up markdown formatting for cleaner text display.
List all available themes with descriptions
Convert text content to a series of images with Xiaohongshu (小红书) style themes. Supported themes: - minimal: Clean white background, good for knowledge/tips content - elegant: Warm beige with serif font, good for novels/essays - warm: Warm gradient with card style, good for lifestyle/emotional content - dark: Dark mode, eye-friendly for night reading Image ratios: - 3:4 (1080x1440): Best for Xiaohongshu feed, takes most screen space - 1:1 (1080x1080): Square format - 4:3 (1080x810): Landscape format
Output schemas completely undocumented. text_to_images and file_to_images return values are not specified, LLMs cannot plan downstream operations or extract image URLs/paths. This violates the fundamental requirement that tools document their output structure.
Parameter descriptions are sparse and non-actionable. Example: 'showCover: Whether to generate a cover page' (40 chars) vs baseline of 72 chars. 'generateAiCover' description is unclear about default behavior and dependency on GEMINI_API_KEY env var, should be explicit in description, not a comment. Lists themes in text_to_images description but not as discoverable output from list_themes.
No error handling guidance. What happens if GEMINI_API_KEY is unset but generateAiCover=true? If outputDir path is invalid? If Playwright browser initialization fails? If text is too long? LLMs need actionable recovery hints like 'Set GEMINI_API_KEY env var and retry' or 'Reduce text length and try estimate_pages first'.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | D | 51 | 2026-07-28+ | v2 |
Destructive/write operations lack confirmation or dry-run pattern. text_to_images is marked WRITE risk and accepts outputDir. file_to_images is marked READ_ONLY but can write to disk (outputDir). No confirmation step, no validation of outputDir path safety (path traversal risk), and no indication of idempotency.
Parameter filePath in file_to_images accepts absolute paths with no validation. Untrusted agent input could trigger path traversal attacks (e.g., '/../../../etc/passwd'). Tool description says it 'Supports .md, .txt, .text files' but does not validate extension server-side, only via description. LLM could pass .sh or .exe files.
charsPerPage parameter is undocumented in terms of valid range, defaults, and effects. Is it per line, per image? What happens if set to 1 or 100000? estimate_pages can take charsPerPage but text_to_images does too, consistency and bounds are unclear.
Weak separation of concerns. list_themes and estimate_pages are utility tools that muddy the tool surface. estimate_pages could be merged into text_to_images as an optional dry-run parameter, and list_themes could be a discovery function. This violates single-responsibility principle and increases cognitive load on LLM tool selection.