MCP server for AI vision analysis via OpenRouter
This Vision MCP server has two tools with partially defined schemas and minimal descriptions. Tool names follow verb_noun convention (analyze_image, list_models), which is good. However, descriptions are generic and lack context about when/why to use tools. The analyze_image tool has a schema with three properties, but descriptions for parameters are perfunctory and lack guidance on valid formats or constraints. The list_models tool has an empty properties object with no parameters, which is reasonable. The server lacks output schema documentation, tool responses are returned as raw JSON strings without structured field definitions. Error handling is present in the McpServer code (try-catch blocks, isError flags) but responses are generic text rather than actionable recovery guidance. No pagination support, no parameter constraints (enums, min/max), and no indication of idempotency. Parameter descriptions do not specify format expectations (e.g., what constitutes a valid file path vs. URL, what OpenRouter model names look like). Overall, this server demonstrates basic tool definition but falls short of production patterns for description clarity, schema completeness, and error guidance.
Analyze an image using AI vision models. Supports file paths and URLs.
Get list of available AI vision models for vision analysis
Parameter descriptions lack actionable format constraints. The 'source' parameter description ('Image source: file path or URL') does not specify valid URL schemes, file path length limits, supported image formats, or maximum file sizes. The 'model' parameter lacks enumeration of valid OpenRouter model names or guidance on discovery. The 'prompt' parameter provides no guidance on minimum/maximum length or expected content structure.
Tool descriptions are generic and do not explain WHEN to use analyze_image vs. alternatives, WHAT data is returned, or ANY prerequisites (e.g., OpenRouter API key setup). Both descriptions lack tokens to help LLMs decide between tools and understand context for selection. Current descriptions read as CLI help rather than LLM-optimized prompts.
Output schema is completely undocumented. Tool responses are returned as JSON-stringified text without schema declaration. LLMs cannot infer response structure, required fields, or data types. The analyze_image response includes 'analysis', 'model', 'timestamp', 'metadata', 'source' fields, but no type information is exposed to the agent.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 50 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 40 | - | v1 |
Error messages are generic text strings with no recovery guidance. When analyze_image fails, the response is 'Error: <message>' with no suggestion for next steps (e.g., 'Invalid file path. Try providing a URL instead.' or 'OpenRouter API error, retry in 30s.'). This leaves the LLM with no actionable path forward.
No pagination, limiting, or result count constraints documented. If list_models returns hundreds of models, no field indicates whether a result is truncated or how many items are available. Large unstructured responses waste tokens and degrade LLM reasoning.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint) declared. Both tools are read-only (no data mutation), but this is not explicitly signaled in the tool definition. Annotations help agents reason about safe retries and composition.
The 'prompt' parameter for analyze_image has no constraints or examples. LLMs may pass multi-paragraph prompts, LaTeX, or structured templates without knowing if the downstream API accepts them. A description like 'Custom analysis prompt (plain text, 10-2000 chars, no markdown or special formats)' would clarify expectations.