MCP server for OpenRouter — chat with 300+ LLMs, analyze/generate images, audio, and video.
This OpenRouter MCP server exhibits strong definition quality across 19 tools with comprehensive descriptions, proper JSON Schema validation, and clear parameter documentation. All tools have non-empty descriptions (50-250 chars typical), input schemas with type definitions, and named parameters. However, there are notable gaps: (1) output schemas are not explicitly documented for any tool, forcing LLMs to infer response structures; (2) several tools lack actionable error guidance (e.g., what to do if a model validation fails); (3) some parameter descriptions reference external URLs without inline constraint documentation. The server demonstrates good adherence to verb_noun naming (chat_completion, analyze_image, generate_video) and includes helpful context like OpenRouter's provider-routing system and caching features. Transport is STDIO-only, which caps the overall protocol readiness score to 50 despite excellent definition quality.
Transcribe or analyze one audio file (WAV, MP3, FLAC, OGG, etc.) with a multimodal model. Output is tagged `_meta.content_is_untrusted: true`.
Analyze one image with a vision model. Accepts a sandboxed local path, https URL, or base64 data URL. Output is model-generated and tagged `_meta.content_is_untrusted: true`.
Describe or analyze one video file (mp4, mpeg, mov, webm). Default model: google/gemini-2.5-flash. Output is tagged `_meta.content_is_untrusted: true`. Large files are fully buffered — prefer short clips.
Send messages to an OpenRouter chat model and get a text reply. Supports provider routing, model suffixes (`:nitro` fastest, `:floor` cheapest, `:free` zero-cost, `:online` web search, `:exacto` tool accuracy), reasoning tokens, web search (`online: true`), and response caching.
Generate audio (speech, music, or sound effects) from a text prompt. Supports multiple audio formats and sample rates.
Output schemas not documented. While input schemas are comprehensive with type definitions, no tool explicitly documents its return value structure (field names, types, nested objects). This forces LLMs to infer response shapes, increasing hallucination risk when planning downstream tool calls or data extraction.
Insufficient error guidance. Tools lack actionable recovery instructions. E.g., validate_model returns 'true/false' but offers no guidance on what to do if false (should the LLM retry, suggest alternatives, or ask the user?). Error responses should classify as retryable vs user-fixable vs fatal.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 66 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 43 | - | v1 |
Generate an image from a text prompt using OpenRouter's image generation models. Supports aspect ratios, sizes, and quality settings. Optional save_path to persist output.
Generate an image using a dedicated image API (not via chat). Returns higher-quality results. Supports resolution, quality, and style settings.
Generate a video from a text prompt. Supported models: Veo 3.1 (Google), Seedance, Wan, RunwayML. Returns job_id for async polling via get_video_status.
Generate a video by animating a static image. Input: local path, https URL, or data URL. Returns job_id for async polling via get_video_status.
Check the status of an async chat completion job started with `start_chat_completion`. Returns the full response when completed, or current status (running/failed) otherwise.
Fetch detailed information about a specific model: pricing, context window, supported features (vision, function-calling, top-k sampling), moderation, and deprecation status.
Poll the status of a video generation job started with generate_video or generate_video_from_image. Returns progress, status (running/completed/failed), and the final video URL when done.
Verify server health and OpenRouter API connectivity. Returns server version, default model, cache status, and API reachability.
Rerank a list of documents by relevance to a query using OpenRouter's reranking models. Returns ranked results with scores.
Search OpenRouter's model catalog by name, capabilities, or provider. Returns paginated results with pricing, context limits, and top-token sampling support.
Transcribe audio to text. Input: local file, https URL, or base64 data URL. Output format options: JSON (default with confidence), plain text, SRT, VTT, or verbose JSON (with tokens/confidence per word).
Start a chat completion as an async background job. Returns a job_id immediately without waiting for the model to respond. Use `get_chat_completion_status` to poll for results. Designed for reasoning models or any request that may exceed MCP timeout limits (~60s).
Convert text to natural-sounding speech using OpenRouter's TTS models. Supports multiple voices and audio formats.
Check if a model exists and is available on OpenRouter. Returns true/false without pricing details.
External URL references without inline constraint documentation. Parameters like `provider` reference 'https://openrouter.ai/docs/features/provider-routing' but do not inline the allowed property names or types. LLMs cannot follow links during inference, constraints must be explicit in descriptions.
Parameter `image_path` in analyze_image documentation includes security warnings ('Bad: "/etc/passwd" (UNSAFE_PATH)') but the tool name and description do not explain the sandboxing mechanism. An LLM reading only the description might not realize paths are restricted.
cache_control and response caching features are mentioned but conditional behavior is unclear. E.g., when `cache: true`, does it always use cached results, or only if a cache hit exists? If cache miss, is a fresh call made? Ambiguity invites LLM errors in assuming availability or staleness.
Async polling pattern (start_chat_completion + get_chat_completion_status, generate_video + get_video_status) lacks timeout and max-attempt guidance. LLMs could poll indefinitely or give up too early. Descriptions should state expected latency ranges and retry limits.
Model selection defaults are underspecified. E.g., 'server default is a free multimodal model' (analyze_image) does not name which model, forcing the LLM to guess behavior. Defaults should be explicit and documented.