MCP server providing tools for fetching arXiv papers and weather forecasts
The server defines 3 tools with basic descriptions and input parameters, but exhibits significant gaps in schema completeness, parameter documentation, and error handling. Tool names follow verb conventions adequately, but descriptions lack LLM optimization and actionable context. Output schemas are entirely undocumented, and parameter descriptions are minimal. No error handling guidance, no permission gates, and no output structure documentation. Per the rubric baseline, 100% of A+ tools have documented return types and all params have descriptions, this server achieves neither. Average across the 3 tools: 42/100.
Fetch summaries and titles of the 20 most recently submitted papers from arXiv related to a given topic. This function queries the arXiv API using the provided topic string and retrieves the 20 latest papers, sorted by submission date in descending order (most recent first). It returns a dictionary where each key is the paper's title and each corresponding value is the paper's abstract (summary).
Get weather alerts for a US state.
Get weather forecast for a location.
Output schemas are completely undocumented. arxiv_papers returns dict[str, str] with no field description; get_alerts and get_forecast return free-text strings with no structure. LLMs cannot plan downstream calls or extract data reliably without knowing what fields to expect.
Parameter descriptions are trivial or missing. 'state' param in get_alerts has description 'Two-letter US state code (e.g. CA, NY)' which includes example values that LLMs tend to reuse literally rather than adapting to context. latitude/longitude in get_forecast lack any format guidance (decimal precision, valid range, coordinate system). Topic param in arxiv_papers has minimal guidance.
Tool descriptions for get_alerts and get_forecast are under 20 characters ('Get weather alerts for a US state.' = 38 chars including period; 'Get weather forecast for a location.' = 36 chars). While technically above 20, they lack context for LLM selection, no information about when to use these vs other weather tools, no mention of geographic scope (US-only for NWS), no guidance on prerequisites or rate limits.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | F | 41 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 39 | - | v1 |
No error handling documentation or recovery guidance. Code catches exceptions and returns None or error strings, but the LLM has no guidance on what to do next. 'Unable to fetch alerts or no alerts found.' is not actionable, should indicate retryability, suggest alternatives, or explain prerequisites.
No input validation constraints documented. Latitude/longitude in get_forecast have no declared range (-90 to +90, -180 to +180). State code in get_alerts has no enum constraint, LLMs will invent invalid codes. arXiv topic has no length limit or character restriction guidance.
arxiv_papers returns up to 20 results with no pagination support. If the result is large, no next_cursor or offset guidance. Tool description states '20 most recently submitted papers' but does not document this as a hard cap or explain pagination behavior.
Tool responses use unstructured strings (get_alerts and get_forecast both return formatted multi-line strings). LLMs must parse natural language output to extract structured data, wasting tokens and increasing error likelihood. Should return arrays of structured objects (e.g., [{event: '...', area: '...', severity: '...'}]) instead.