Providing tools to LLMs through Model Context Protocol (MCP)
Server provides 4 tools with basic naming and schema coverage. Tool names follow verb_noun pattern (web_search, news_search, image_search, video_search), which is good. However, descriptions are generic and repetitive across all tools. Parameter descriptions are minimal, 'A detailed search query string' and 'Number of results to return (default: 5)' lack context about constraints, format, or when to use each tool. Output schemas are not documented, callers cannot predict what fields to expect in the response (title, snippet, URL are mentioned in docstring but not in a formal schema). All tools are read-only (appropriate risk classification), but there is no guidance on error handling, recovery steps, or result limits. The tools are simple wrappers around DuckDuckGo with identical structure across all four, which increases the risk of LLM confusion when selecting between them.
Search the web for images for the given query. Provide a detailed query for the best results, including keywords, locations, etc.
Search the web for news on a given query. Provide a detailed query for the best results, including keywords, locations, etc.
Search the web for videos for the given query. Provide a detailed query for the best results, including keywords, locations, etc.
Search the web for information on a given query. Provide a detailed query for the best results, including keywords, locations, etc.
Output schemas are not documented. Tools return list but do not declare field structure (title, snippet, url, etc.), LLMs cannot predict response format or plan downstream operations.
Tool descriptions are generic and nearly identical across all four tools. Each repeats 'Provide a detailed query for the best results, including keywords, locations, etc.', no distinction of when to use news_search vs web_search vs image_search. LLMs will struggle to select the right tool.
Parameter descriptions lack constraint details. 'Number of results to return (default: 5)' does not specify minimum/maximum bounds. What happens if LLM passes 1000? Is there a cap? What is the performance impact? Unbounded numeric parameters invite absurd values.
Inferred effective spec: <=2025-11-25.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 61 | <=2025-11-25 | v2 |
| 2026-03-09 | F | 49 | - | v1 |
No error handling or recovery guidance. Tool descriptions do not tell LLMs what to do if search fails, times out, or returns no results. No guidance on retryable vs fatal errors.
No result pagination or limit enforcement documented. Tools accept 'num_results' but do not state a maximum. Large result sets can blow context window. No mention of pagination in return schema.