An AI-powered news aggregation and notification system that fetches trending news, manages user subscriptions, and sends email notifications for trending topics
Single tool 'search_hot_news' has moderate structural definition but significant gaps. Tool name is action-verb correct (search_), description is present but generic (123 chars, below quality baseline of 194 chars average). Input schema is visible with types and defaults, but parameter descriptions are thin (7-24 chars each). Output schema is entirely undocumented, callers cannot know what fields to expect from the returned articles list. No error handling guidance provided. Tool modifies no state (read-only) but lacks documentation of what the returned 'articles' object structure contains.
Search for hot/trending news given query using SerpAPI
Output schema not documented. Tool returns articles list with fields like 'id', 'source', 'url', 'title', 'content', 'published_at', 'language', 'recency_score', 'source_authority', 'engagement_keywords', 'is_breaking', 'search_query', 'search_timestamp', but no schema definition provided. LLMs cannot infer what fields are present or their types, forcing guesswork about how to parse results.
Parameter descriptions too brief and lack detail. 'query' described as 'Search query' (12 chars), 'language' as 'Language code' (13 chars), 'timeframe' as 'Time filter: "1h", "24h", "7d", "1m", "1y" or custom date' (56 chars). While timeframe includes enum values via text, best practice is to formalize as enum constraint in schema, not text examples. Descriptions should be 50-150 chars with context on when to use each parameter.
Tool description under quality baseline. 'Search for hot/trending news given query using SerpAPI' is only 62 chars; production baseline is 194 chars average (p90=392). Description lacks WHEN to call this vs alternatives, WHAT the output structure contains, and HOW to interpret scoring fields (recency_score, source_authority). LLM cannot determine selection criteria.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 47 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 35 | - | v1 |
No error handling guidance. Tool catches exceptions silently and returns empty list [] on error. LLM receives no indication of what went wrong (API key missing, rate limit, network failure, malformed query) and cannot self-correct or provide user feedback. SerpAPI failures should return structured errors with recovery hints.
Timeframe parameter values documented as text examples, not enum. Code includes tbs_map dict with hardcoded values ('1h', '24h', '7d', '1m', '1y'), but schema does not enforce this as an enum with description stating 'Must be one of: 1h, 24h, 7d, 1m, 1y, or a custom date'. Free-form strings invite hallucinated invalid values.
num_results parameter lacks bounds documentation. Default is 5, but no min/max stated. SerpAPI may accept 1-100 or have other limits, the tool description should specify valid range (e.g., '1 - 100') to prevent LLM from requesting unreasonable result counts.
Helper functions (calculate_recency_score, get_source_authority, extract_hot_keywords, is_breaking_news) referenced in code but extract_hot_keywords and is_breaking_news are not shown. Partial source code visibility, cannot fully verify output field definitions.