MCP server for searching AWS Open Data Registry datasets
This server demonstrates good definition quality with properly structured tool schemas, clear naming conventions, and reasonable descriptions. Both tools follow verb-noun patterns and include proper input/output schemas with Zod validation. Descriptions are adequate but could be more action-oriented. The main gaps are: (1) output schemas lack pagination/limit enforcement guidance despite returning potentially large datasets, (2) error handling descriptions are minimal, (3) no tool annotations for read-only hints despite all tools being safe. The search_datasets tool's 'detail' parameter is well-designed with enum constraints, and parameter descriptions are generally solid. Composition is clean, two single-responsibility tools with clear data flow.
Get detailed information about a specific dataset by its ID (e.g., 'sentinel-1')
Search for datasets in the AWS Open Data Registry. If no query is provided, lists all datasets. Returns datasets matching the search query in their name, description, or tags.
Output schema lacks pagination or result-size enforcement guidance. search_datasets returns an unbounded array of datasets but provides no next_cursor, total_count, or page_number. With 500+ datasets in AWS Open Data Registry, returning all matches could blow context windows. Description mentions 'limit' parameter but schema/implementation lacks a 'total' or 'next_cursor' field.
Missing tool annotations. Both tools are read-only and safe to retry/retry, but lack readOnlyHint or idempotentHint annotations. Modern MCP clients and agents use these hints to optimize caching and retry behavior. Schema should include 'readOnlyHint': true in tool metadata.
Error handling descriptions are absent. Neither tool documents what happens if: (1) the query times out downloading the registry, (2) a dataset YAML is malformed, (3) get_dataset is called with a non-existent ID. These are production failure modes, the LLM needs recovery guidance (e.g., 'try search_datasets to find the correct ID'). Code logs errors but does not return actionable messages.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | B | 78 | 2025-06-18+ | v2 |
| 2026-03-09 | D | 55 | - | v1 |
get_dataset description lacks clarity on failure modes and the expected format of the 'id' parameter. It states 'e.g. sentinel-1' but does not make it clear whether IDs can be looked up via search_datasets or if users must know them in advance. The description should hint at the relationship: 'Use search_datasets to find IDs if unknown.'
Response field naming could be more consistent for agent chaining. search_datasets returns results with 'Name' (capitalized) and 'id' (lowercase); get_dataset returns 'dataset' wrapping the full object. An agent chaining these calls must handle inconsistent casing and nesting, this adds cognitive overhead. Standardize to 'name' and 'id' across all outputs.