Unified async web search across multiple backends with MCP server, CLI, and admin UI
The server defines 4 tools with clear action-verb names (search, search_news, get_answer, get_provider_info) that follow naming conventions. All tools have descriptions (ranging 56-148 chars) and documented input schemas with types. However, several critical gaps prevent a higher score: (1) Parameter descriptions are minimal or missing for max_results constraints (no mention of 1-50 range for search tool, or validation logic); (2) Output schemas are documented only in docstrings, not in structured form visible to the schema evaluator, responses are described as 'List of results with...' but no JSON Schema output definition is provided; (3) No error handling guidance, no recovery patterns, no actionable error messages, no guidance for LLM on what to do if a search fails; (4) Tool descriptions lack WHEN/WHY context for selection, e.g., when should an agent call search vs search_news vs get_answer?; (5) No documentation of rate limits, timeouts, or dependencies between tools. The tools are functionally sound and well-named, but lack the signal-rich descriptions and error patterns expected of production-grade tools. Average per-tool score: 62.
Get a direct answer to a question. Args: query: The question to answer. Returns: Dict with 'answer' (string or null) and 'provider' name.
Get information about the current search provider. Returns: Dict with name, configured, api_key_set, features, rate_limit_remaining.
Search the web. Args: query: The search query string. max_results: Maximum number of results (1-50, default 10). Returns: List of results with title, url, snippet, score.
Search for recent news articles. Args: query: The search query string. max_results: Maximum number of results (default 5). Returns: List of news results with title, url, snippet.
Output schemas not formalized. All tool responses are documented in docstrings (e.g., 'List of results with title, url, snippet, score') but no JSON Schema output definitions are provided. This forces LLMs to infer output structure, increasing hallucination risk and preventing downstream tool chains from validating inputs. See pattern:tool and pattern:response-shaper.
Parameter constraints missing from descriptions. The 'search' tool accepts max_results 1-50 per the rubric description, but this constraint is NOT documented in the parameter description, only visible in schema (if at all). The 'max_results' param descriptions say 'Maximum number of results (1-50, default 10)' and '(default 5)' but do NOT specify the valid range in human-readable form. Per pattern:constrained-input, constraints should be explicit in descriptions so LLMs understand them.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 64 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 47 | - | v1 |
No error handling guidance or recovery patterns. Tools have no documented error responses, no guidance on what to do if a search returns 0 results, no timeout behavior, no rate-limit handling. Per pattern:recovery-guide and pattern:error-classification, error responses must tell the LLM what to do next (e.g., 'If no results found, try a broader query or call get_answer').
Tool selection context missing. Descriptions do not explain WHEN to use each tool. For example, should an agent call 'search' for 'What is the weather in NYC?' or 'get_answer'? Should it use 'search_news' for 'Recent news on AI' or 'search'? Lack of selection context forces LLMs to guess, causing tool misuse. Per pattern:tool-description, descriptions must answer 'What does it do? WHEN should the LLM call it instead of a similar tool?'
No tool-chaining metadata. Response fields do not include IDs or references that might be needed by downstream tools. For example, 'search' returns title, url, snippet, score, but if a hypothetical 'fetch_article' or 'analyze_content' tool existed, it would need a stable reference (url suffices here, but this is not documented as a chaining contract). Per pattern:tool-chain, ensure tool A's output contains IDs tool B needs.
get_provider_info returns sensitive state without context. The output includes 'api_key_set' (boolean) and 'rate_limit_remaining', but no guidance on what the agent should do if api_key_set=false or rate_limit_remaining=0. Should it fail? Retry? Inform the user? Per pattern:tool-description, descriptions must include prerequisites and dependency hints.