Multi-agent research pipeline that decomposes queries into sub-questions, searches the web via Tavily, scores source credibility, and synthesizes comprehensive markdown reports with findings, knowledge gaps, and cited sources.
Single tool with well-structured schema and comprehensive description. Input validation via Pydantic with proper constraints (min/max length, enum). Description is detailed (380+ chars) and explains the three-stage pipeline, caching behavior, and output format. Parameters are typed and described. However, output schema is not formally documented, only described in prose. Error handling is present but generic (catches all exceptions, returns truncated messages). No tool annotations (readOnlyHint, idempotentHint). Naming is clear (verb_noun: deep_researcher_research) but slightly redundant.
Run a multi-agent research pipeline that decomposes a query into sub-questions, searches the web in parallel via Tavily, scores source credibility, and synthesizes a comprehensive markdown report with findings, knowledge gaps, and cited sources. The pipeline has three stages: 1. Planner — breaks the query into 3-5 targeted sub-questions (GPT-4o) 2. Searcher — runs parallel Tavily web searches with credibility scoring 3. Synthesizer — produces a structured markdown report (GPT-4o). Results are cached for 24 hours. Identical (query + context + depth) combinations return instantly on subsequent calls.
Output schema not formally documented. Description states 'Markdown research report with executive summary, key findings per sub-question, knowledge gaps, and numbered source citations' but no structured schema (JSON Schema) is provided. LLMs cannot plan downstream tool calls or extract fields reliably.
Error handling is generic and truncates messages to 200 chars. Errors like 'Research pipeline failed, ValueError: ...' do not guide the LLM on recovery steps (retry, adjust query, check API keys). Should return actionable guidance per error type.
No tool annotations (readOnlyHint, idempotentHint, destructiveHint). The tool is read-only and idempotent (caches results for 24h), but these hints are not declared in the schema. Clients cannot infer safety properties.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 69 | 2026-07-28+ | v2 |
Tool name 'deep_researcher_research' is slightly redundant (verb_noun where noun repeats the server name). Consider 'research' or 'run_research' for clarity and brevity.
No pagination or result limiting documented. If markdown report grows very large (e.g., 50+ sources), it could exceed context windows. Should document max output size or offer a summary mode.