The server registers 7 tools with complete input schemas and descriptions. All tools follow the verb_noun naming pattern (list_, get_, search_, compare_, refresh_). Descriptions are present and adequately detailed (ranging 40-200 chars). However, parameter descriptions lack specificity around constraints, ranges, and error guidance. Output schemas are not documented, LLMs cannot infer what fields to expect from responses. No error handling guidance is provided. The implementation uses client-side caching with a 5-minute TTL, but tools do not document this behavior or guide recovery when cache is stale. All tools are read-only (no destructive operations), which is positive for safety.
Tools (7)
compare_modelsread onlysource verified72/100
Compare multiple models side by side
get_cheapest_modelsread onlysource verified75/100
Get the cheapest models meeting specified criteria
get_modelread onlysource verified73/100
Get detailed information about a specific model
list_modelsread onlysource verified85/100
List available models with optional filtering and pagination. Returns max 50 models by default to stay within token limits.
No output schemas documented. Tools return unstructured data with no field-level type hints. LLMs cannot infer what fields to expect, how to chain tools, or what IDs to extract for downstream calls.
Parameter descriptions lack actionable constraints. E.g., 'sort_by' enum lists field names but doesn't explain the sort behavior; 'limit' has no min/max guidance; 'max_price_per_1m_tokens' has no range. LLMs cannot validate inputs without explicit bounds.
No error handling guidance. Tools have no documented error responses, recovery paths, or what to do if a model_id is invalid or API times out. LLMs get raw errors with no context.
list_modelsget_modelcompare_modelssearch_models
Recommendations
Document output schemas for all 7 tools. E.g., list_models should return { models: [ { id: string, provider: string, name: string, pricing: { input: number, output: number }, context_length: number } ], total_count: number, has_more: boolean }. This enables LLMs to chain tools (e.g., extract model IDs from list_models and pass to compare_models).
Add min/max constraints to numeric parameters. E.g., limit: 'Maximum number of models to return (1 - 200, default: 50)'; min_context_length: 'Minimum context length (e.g., 4000, 8000, 128000)'.
Expand parameter descriptions with format/constraint hints. E.g., modality: 'Input modality (text|image|audio|video); omit to return all modalities'. sort_by: 'Field to sort by: price (ascending = cheapest first), context_length, name (alphabetical), created (newest first)'.
Add error handling guidance to each tool description or via a separate error documentation. E.g., 'If a model_id is not found, suggests available models with similar names. If API rate-limited, returns 429 with retry-after guidance.'
Clarify cache behavior in tool descriptions. E.g., list_models: '...Returns models from a local 5-minute cache (updated every 5 min automatically). Call refresh_cache to force an immediate update.' get_model: '...Note: single model queries do not bypass cache.'
Expand get_model and list_providers descriptions to explain their relationship to list_models. E.g., get_model: 'Retrieve full details for a single model (name, pricing, context length, modalities, capabilities). Use after list_models to get more info about a specific model.' list_providers: 'Discover all available providers at a glance. Call this first if the user asks for a list of AI companies.'
Cache behavior undocumented in tool descriptions. Tools silently use a 5-minute cache, but there's no hint to LLMs that refresh_cache exists or when calling it is necessary. The relationship between list_models and refresh_cache is opaque.
Tool descriptions are minimal (40 - 65 chars) and lack context about when to use each tool. E.g., 'List all unique providers with model counts' doesn't explain when an LLM should call this instead of list_models.
get_modellist_providersrefresh_cache
Add discovery hints to tool descriptions. E.g., search_models: 'Use this to find models by name or description. If the user says 'I want a model for video', call this with query='video' first, then compare_models if narrowing down.'
Return structured per-item errors for batch operations. If compare_models is called with invalid model IDs, return { compared: [...], errors: [ { model_id: 'invalid', reason: 'not found', did_you_mean: ['claude-3-opus', 'gpt-4'] } ] } instead of failing entirely.
Consider adding a batch version: list_models_batch(limit=1 - 200, queries=[...]) to support multi-criterion searches in one call (e.g., find all models with context_length >= 100k AND price <= $0.001 AND modality='text'). Current design requires multiple sequential calls.