MCP server providing DuckDuckGo web search via HTTP with SSE (Server-Sent Events)
The DuckDuckGo MCP server has one tool (web_search) with a complete JSON Schema and clear naming. The tool description is adequate (72 chars, baseline 194 avg). All five parameters (query, max_results, all_results, region, safesearch, timelimit) have type declarations and descriptions. The schema enforces constraints via enum, minimum/maximum, and defaults. However, the output schema is NOT documented in the visible source code, the response structure from the DuckDuckGo API is wrapped but not formally declared for the LLM. Error handling guidance is absent (e.g., what to do if DuckDuckGo is unreachable, or if a query returns zero results). The tool does not include composition hints (e.g., when to use region vs timelimit). Overall: solid parameter definitions and naming, but incomplete output documentation and error guidance prevent a higher score.
Search the web using DuckDuckGo
Output schema not documented. The tool description does not specify what fields the DuckDuckGo response contains (e.g., title, URL, snippet, rank). LLMs cannot plan downstream tool calls without knowing the output structure.
No error recovery guidance. If DuckDuckGo is unreachable or the query yields zero results, the tool should return a structured error with actionable next steps (e.g., 'Try a simpler query', 'Check your internet connection'). Current error handling logs stack traces but provides no LLM-actionable guidance.
Tool description is brief (72 chars, below baseline 194 avg). It does not explain WHEN to use web_search (vs. alternatives) or clarify that results are capped at 10 items. A more complete description would state: 'Search the web using DuckDuckGo. Returns up to 10 results ranked by relevance. Use region and timelimit parameters to narrow results geographically or by recency.'
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | B | 71 | <=2025-11-25 | v2 |
Parameter description for 'query' lacks guidance on format and length. It should specify: 'Natural language search term (1 - 200 chars, no special operators required).' LLMs may otherwise pass excessively long or malformed queries.
Redundant parameters 'max_results' and 'all_results' create ambiguity. When all_results=true, does max_results apply? The descriptions do not explain the mutual relationship. Simplify: keep max_results (1 - 10) as the sole control, remove all_results.
No output pagination or limit enforcement in the tool interface. If the LLM does not set max_results or all_results, what is the default behavior? The description says 'default: 5', but there is no explicit per-result-item schema. Without documenting per-item fields (rank, title, URL, snippet), LLMs cannot reliably extract and chain results.