Mixed quality across 10 tools. Strengths: all tools have descriptions (10-150 chars), input schemas with types and constraints (enums, min/max bounds), and clear verb-noun naming. Weaknesses: descriptions are terse (many under 50 chars), output schemas are not documented, no error handling guidance, no tool annotations (readOnlyHint/destructiveHint). Parameters lack contextual hints for chaining and recovery. The Apify tools (run_actor, get_actor_info, etc.) are well-constrained but lack operational guidance; DuckDuckGo tools are similarly sparse. No tools declare recovery paths on failure or guide multi-step composition.
No output schemas documented. Tools return results but LLMs cannot see what fields to expect. Breaks tool chaining and forces agents to infer structure from calls.
Descriptions are terse (avg ~60 chars for DuckDuckGo tools, ~80 for Apify). Lack operational context: when to use this tool vs alternatives, what happens on error, prerequisites. Rubric requires 10 - 1024 chars; these are at lower end and omit guidance.
No error handling guidance. Tools define parameters but not recovery paths. E.g., run_actor may timeout (max 3600s), but no guidance on what to do on timeout.
Recommendations
ADD output schemas to all tools. Document the response structure: which fields are returned, their types, and which are required. E.g., run_actor should document {run_id: string, status: 'RUNNING'|'SUCCEEDED'|'FAILED', started_at: ISO8601, ended_at?: ISO8601, ...}. This is critical for tool chaining.
EXPAND descriptions to 100 - 150 chars. Current descriptions like 'Get detailed information about an Apify Actor' lack context. Improve to: 'Retrieve actor metadata including name, category, pricing, and latest build info. Use this to validate actor availability before calling run_actor or to understand input requirements.'
ADD tool annotations. Mark run_actor with destructiveHint=true (creates a run, consumes credits). Mark search_*, get_* with readOnlyHint=true. This helps agents reason about plan safety.
ADD error handling guidance in descriptions. E.g., run_actor: 'If timeout occurs, use get_run_status(run_id) to check progress. Retries are safe (idempotent with same actor_id + input). Note: each retry incurs Apify credits.'
ADD parameter context hints. E.g., 'query' parameter in search_news: 'Text search query (max 400 chars). Return includes URL, title, snippet, use fetch_content(url) to retrieve full article.'
DOCUMENT pagination and result limits. E.g., search_images: 'Returns up to max_results items (capped at 50). Images include URL, source, and dimensions. No cursor/offset pagination, retry with refined query for different results.'
ADD examples to descriptions (not parameter names). E.g., search: 'Search for news, articles, or web results. Example: "Python 3.13 release". Returns results sorted by relevance with snippets.' This guides LLM selection without hard-coding values.
No tool annotations (readOnlyHint, destructiveHint, idempotentHint). run_actor is destructive (WRITE risk), but no annotation. search_* and get_* are idempotent READ-only, but unmarked. Per spec 2026-07-28, annotations help agents plan safely.
Parameter dependencies undocumented. E.g., get_dataset_items has format enum (json/csv/xml) and clean flag, but no guidance on which formats support clean or what clean means for each format.
No chaining IDs in responses. E.g., run_actor triggers a run with run_id, but response schema unknown, agent cannot infer whether to pass run_id to get_run_status immediately.
Generic parameter descriptions in DuckDuckGo tools. 'The search query' (search/search_news/search_images) omits format, max length (stated as 400 in schema but not description), or examples of invalid input.
searchsearch_newssearch_images
ADD parameter dependencies. E.g., scrape_url: 'extract_data object fields (title, content, links, images) apply only when javascript_enabled=true. Set to false for static HTML parsing.'
DOCUMENT what to do on failure. E.g., scrape_url: 'If JavaScript rendering fails, retry with javascript_enabled=false. If no results, broaden selectors in extract_data. Max 10 URLs per call, split larger batches.'
SPECIFY identity resolution. E.g., if Apify tools accept 'actor_id' as name OR UUID, document: 'actor_id can be a UUID (e.g., "xyz...") or name (e.g., "apify/web-scraper"). Names are resolved to UUID automatically.'
ADD per-item result status for batch operations. E.g., if get_dataset_items returns a list, include per-item flags: {success: true, data: {...}} or {success: false, error: 'field validation failed'}. This avoids blanket errors on partial failures.
REPLACE enums with descriptions. E.g., search_actors category enum, add: 'Category filters results to a specific domain. E.g., E_COMMERCE for web scraping actors, SOCIAL_MEDIA for Instagram/TikTok tools.' This contextualizes the choice for LLMs.