Two tools with clear action-verb naming (outsource_text, outsource_image) and reasonable parameter schemas. However, several definition gaps reduce overall quality: (1) Descriptions lack WHEN-to-use guidance and prerequisites despite being present; (2) Parameter descriptions exist but are minimal (provider and model descriptions are generic strings with example lists rather than formal constraints); (3) No documented output schema beyond 'str' return type; (4) Error handling returns plain error strings rather than structured recovery guidance; (5) Tool descriptions do not clarify what happens on failure or what the agent should do next; (6) No pagination, batch, or composition patterns despite delegating to external LLMs.
Delegate image generation to an external AI model. Use this when you need to create visual content.
Delegate text generation to another AI model. Use this when you need capabilities or perspectives from a different model than yourself.
Parameter enum constraints missing. 'provider' param accepts free-form strings with example list in description (e.g., 'openai', 'anthropic', 'google', 'groq') rather than declared enum. This invites hallucinated provider names not in PROVIDER_MODEL_MAP. For outsource_text, valid providers are: openai, anthropic, google, groq, deepseek, xai, perplexity, cohere, fireworks, huggingface, mistral, nvidia, ollama, openrouter, sambanova, together, litellm, vercel, aws, azure, cerebras, meta, deepinfra, ibm. These should be an enum constraint, not examples in description text.
No documented output schema. Both tools return plain strings ('str'). outsource_text returns 'The text response from the external model, or an error message if the request fails', but there is no schema structure defining whether an error is distinguishable from success (e.g., no error_code field, no success boolean). outsource_image returns 'The URL of the generated image' but does not document format (is it guaranteed a URL string, or could it be an error message?). Agents cannot distinguish success from failure or plan downstream steps reliably.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 57 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 47 | - | v1 |
Error handling returns unstructured error strings instead of actionable recovery guidance. Both tools catch exceptions and return 'Error <message>: <exception>' as a plain string. Per the rubric, error responses must tell the LLM what to do next, e.g., 'Provider unknown. Try one of: openai, anthropic, google' or 'Model not found. For GPT models, try: gpt-4o, gpt-4-turbo, gpt-3.5-turbo'. Current implementation gives no recovery path.
outsource_image description states 'Currently only OpenAI is supported' but server imports and exposes 20+ providers in PROVIDER_MODEL_MAP. This is misleading, the tool hardcodes openai-only support, but a user reading the broader context might assume other providers work. Clarify that while the server supports multiple text generation providers, image generation is gated to OpenAI only.
Parameter descriptions contain example values rather than formal constraints. E.g., 'The specific model identifier (e.g., "gpt-4o", "claude-3-5-sonnet-20241022", "gemini-2.0-flash-exp")', LLMs latch onto example values and may pass them literally in edge cases. Replace with a pattern or enum, or state 'Use the model ID from your provider's API docs (e.g., gpt-4o for OpenAI, claude-3-5-sonnet-20241022 for Anthropic).'
No documented failure modes or rate limits. outsource_text and outsource_image call external LLM APIs without documented timeouts, retry behavior, or rate-limit handling. An agent could hit OpenAI rate limits or timeout hanging on a slow model and get stuck. Tool descriptions should note: 'This tool makes synchronous calls to external APIs; expect latencies of 5-60s depending on model. If the call times out, the LLM provider may be overloaded, consider retrying with a different model.'
Tool descriptions lack WHEN-to-use guidance. 'Delegate text generation to another AI model. Use this when you need capabilities or perspectives from a different model than yourself.' is accurate but vague. Better: 'Use when the current model lacks specialized knowledge (e.g., call with provider=deepseek, model=deepseek-coder to get code optimization advice). Or use with a different large model for a second opinion on reasoning tasks. Do not use for simple factual queries the current model can answer.'