Give text-only LLM coding agents (DeepSeek V4 Flash/Pro, Qwen, Kimi) vision via any multimodal model. One analyze_image tool — paste or drag in one or many images, analyzed in a single call, scoped to the current session. Works with Claude Code, opencode, Codex, Kimi Code, PI.
Single-tool server with well-structured Zod schema and comprehensive, multi-line description. The tool name 'analyze_image' is appropriately action-oriented and descriptive. Input schema includes proper type definitions (union of string/array with constraints) and per-parameter descriptions. The description is unusually comprehensive (1,000+ chars) and guides the agent through all valid image sources and task options. However, output schema is not formally documented in the visible code, the tool handler returns results but the expected structure is not declared in the registration or documentation. Error handling and recovery guidance are absent from the tool definition itself. The server implements a sophisticated image caching system and multi-provider abstraction internally, but these implementation details are not exposed to the agent in a way that would influence tool selection or error recovery.
Analyze an image using a multimodal model and return a detailed text description. The vision model sees the image; the calling agent is text-only and cannot. Sources for `image` (pick one, or pass an array to analyze several at once): - "path": absolute or relative path to a local image file (PNG/JPEG/WEBP/GIF) - URL: http(s) URL to an image on the web or a local server - "data:...": base64 data URI, e.g. data:image/png;base64,<payload> - "clipboard": read the image currently copied to the system clipboard - "recent": auto-find the most recently pasted image in THIS session - "session": auto-find EVERY image pasted in this session (analyze them all in one call) - "raw": the string itself is the literal raw image bytes - ["a.png", "b.png", ...]: array of any of the above — all analyzed in one request Pick `task` for common jobs (describe | ocr | ui | layout | qa) or pass your own `prompt`. `detail` defaults to "high" for maximum completeness. Use `save_to` to write a long description to a file and get back only a path + summary.
Output schema not documented. The tool handler returns results but the expected response structure (fields, types, format) is not declared. Agents cannot plan downstream operations without knowing what fields to expect.
No error handling guidance in tool description. Tool can fail for missing files, invalid URLs, unsupported image formats, API rate limits, or provider configuration issues. Agent receives raw errors with no recovery hint.
Tool description embeds example values ('a.png', 'b.png') which LLMs may reuse literally. Should use enums or regex patterns instead, or rely on test examples separate from descriptions.
Parameter 'detail' enum has 'high' as implied default but no explicit default value in Zod schema. Optional enum parameters should state the default in the description or Zod.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 65 | 2026-07-28+ | v2 |