MCP server for AI model lifecycle data — check whether a model id is deprecated or retiring, and what to migrate to, across OpenAI, Anthropic, Azure, AWS Bedrock, Google, and Cohere.
The server defines 5 tools with well-structured schemas and detailed descriptions. All tools use action verbs (check_, list_, suggest_, search_, recent_) and have input schemas with proper type definitions. Descriptions are thorough and context-aware, explaining WHEN and WHY to use each tool. Error handling is explicit with isError flags and actionable recovery guidance. The main weakness is lack of explicit output schema documentation, tools return formatted text without documented structure, forcing LLMs to parse unstructured output. Additionally, descriptions include example values (e.g., 'gpt-4o', 'claude-sonnet-4-5-20250929') which could cause literal reuse, though in this context (model IDs) that is less problematic. All tools declare readOnlyHint=true, signaling safety to hosts. Parameters are well-constrained (enums, ranges, defaults). Tool composition is clean: each does one thing (check status, list retiring, suggest replacement, search, track changes). No security issues detected, all tools are read-only catalog lookups with no state mutation or credential exposure.
Check whether an LLM model id is safe to use, deprecated, or retired, and what to migrate to. Accepts the exact string used in code (e.g. 'gpt-4o', 'claude-sonnet-4-5-20250929', 'anthropic/claude-opus-4-1'). Call this before writing or changing any hardcoded model id.
List tracked models scheduled to retire within a time window, soonest first. Use this to audit a codebase or plan migration work.
Lifecycle state changes across all tracked providers, newest first. Use this to answer 'what model deprecations happened recently?'.
List tracked models, optionally filtered by provider and lifecycle state. Use this to discover what is currently available from a provider.
Given a deprecated or retiring model, return the provider's recommended replacement(s) and, when none is published, active models from the same provider to consider.
Output schemas not explicitly documented. Tools return text via text() or failure() wrappers, but the structure of that text (paragraphs, sections, formatting rules) is not formalized. LLMs must infer structure from formatting conventions alone.
Descriptions include example values (e.g., 'gpt-4o', 'claude-sonnet-4-5-20250929', 'anthropic/claude-opus-4-1') which violate the pattern of avoiding sample values in descriptions. While model IDs are unlikely to be reused literally in wrong contexts, this sets a precedent and could confuse LLMs if they treat examples as preferred formats.
No pagination or limit enforcement visible in list_retiring_models and search_models. The code calls client.listModels with limit:500, but the tool parameters accept withinDays and limit without documented maximums to prevent token exhaustion. search_models has a limit param but no documented cap.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | B | 75 | 2026-07-28+ | v2 |
Tool descriptions do not explicitly state dependencies or call ordering. For example, check_model returns 'not found' for unknown models, but describe_error handles network failures, the LLM cannot distinguish between 'model truly absent' and 'network unreachable' without reading implementation details.