A weather MCP server that wraps OpenWeatherMap API as MCP tools, integrated with a LangChain agent backend and Streamlit frontend for natural-language weather queries.
Three tools with clear verb-noun naming (geocode_city, get_current_weather, get_forecast) and basic descriptions. All tools have input schemas with type definitions and parameter descriptions. However, descriptions are minimal (34 - 60 chars, below the 50 - 200 char baseline for LLM optimization), lack context on when to use each tool, and omit dependency hints. Output schemas are not documented. Error handling returns structured dicts but lacks recovery guidance or error classification. No tool annotations (readOnlyHint, idempotentHint). Rate limiting and input validation are implemented server-side but not surfaced in tool descriptions.
Resolve a city name to latitude and longitude coordinates.
Get current weather conditions for a city.
Get a daily weather forecast for a city (1-5 days).
Descriptions are too brief (34 - 60 chars vs. 50 - 200 char baseline). Lack context on when to use each tool, dependencies, and what data is returned. LLMs cannot reliably select the right tool or understand output structure.
Output schemas are not documented. LLMs do not know what fields to expect from each tool, forcing them to guess at response structure and complicating downstream tool chaining.
Error responses return structured dicts but lack recovery guidance. E.g., 'Rate limit exceeded' does not tell the LLM whether to retry, wait, or ask the user. No error classification (retryable vs. user-fixable vs. fatal).
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | C | 69 | 2026-07-28+ | v2 |
No tool annotations (readOnlyHint, idempotentHint). All three tools are read-only and idempotent, but this is not declared in the schema. Agents cannot infer safety properties without explicit hints.
Parameter 'days' in get_forecast has a default (3) and description mentioning '1-5 days', but no explicit min/max constraints in the schema. LLMs may pass invalid values (0, 6, 100) without validation feedback.