Use any LLMs (Large Language Models) for Deep Research. Support SSE API and MCP server.
The server defines 5 tools with explicit schemas and descriptions visible in src/app/api/mcp/server.ts and package.json. All tools have descriptions and input schemas with typed parameters. However, there are significant gaps: (1) Tool names lack clear action verbs, 'deep-research', 'write-research-plan', 'generate-SERP-query' do not follow verb_noun convention; (2) Several descriptions are vague about what the tool returns and when to use it vs alternatives; (3) Output schemas are not documented, callers don't know the response structure; (4) Parameter descriptions vary in quality; (5) No error handling guidance or recovery paths documented. The schema quality is moderate, parameters are typed but some descriptions are minimal.
Start deep research on any question, obtain and organize information through search engines, and generate research report.
Generate a list of data collection tasks based on the research plan.
Generate SERP queries based on the research plan.
Write a final research report based on the research plan and the results of the information collection tasks.
Generate research plan based on user query.
Tool names do not follow verb_noun convention. 'deep-research', 'write-research-plan', 'generate-SERP-query' lack clarity. LLMs cannot infer intent from these names alone. Should be prefixed with action verbs like 'start_deep_research', 'create_research_plan', 'generate_search_queries'.
No output schemas documented. Callers cannot know what fields the tools return or their types. LLMs must guess at response structure, risking failed downstream chaining. Each tool description should include 'Returns: {...}' schema.
No error handling or recovery guidance. Tools do not document what errors can occur, when they are retryable, or what the LLM should do next. E.g., what if search fails? What if an API is rate-limited?
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 63 | 2026-07-28+ | v2 |
| 2026-03-09 | C | 69 | - | v1 |
Vague tool descriptions. 'Generate a list of data collection tasks based on the research plan' (generate-SERP-query) does not explain WHEN to call this tool vs alternatives, WHAT it returns, or HOW it differs from search-task. Descriptions should be 50-200 chars and answer: what, when, why.
Parameter 'query' appears in multiple tools (deep-research, write-research-plan, generate-SERP-query, search-task) with identical or near-identical descriptions. No distinction of what each tool expects or produces. This creates ambiguity for LLM tool selection.
'generate-SERP-query' name uses non-standard casing (mixed-case with hyphen and SERP acronym). Should be 'generate_serp_queries' or 'generate_search_queries' to match verb_noun convention and lowercase/underscore standard.