MCP server for Autario | search, query, join, and analyze 2,700+ public datasets (World Bank, FRED, Eurostat, OECD, WHO, ECB, US Census, IMF). Cross-dataset joins via shared ontology, statistical analysis (correlation, regression, drivers, lag), and chart publishing. Works with Claude, ChatGPT, Cursor, and any MCP-compatible AI.
Autario MCP presents a well-structured data query tool suite with good naming conventions, clear descriptions, and detailed input schemas. All 4 tools follow verb_noun naming (search_, get_, query_) which aids LLM intent detection. Descriptions are comprehensive (180-380 chars) and include usage guidance. Input schemas are fully specified with types, required fields, and parameter descriptions. However, output schemas are not formally documented, responses are inferred from API behavior rather than declared in the tool definition. Error handling is minimal (tool code shows no specific error recovery guidance for LLMs). Tool composition is sound (each does one thing), though the aggregate query feature in query_dataset borders on scope creep. No security issues detected (no secrets in params, readonly tools). The server lacks progressive complexity features like pagination guidance in descriptions and does not explain when to use list_indicators vs search_datasets.
Get full metadata for a specific dataset including title, description, publisher, category, keywords, row count, and creation date.
Get the column names, data types, and total row count for a dataset. Always call this before query_dataset to understand the available columns for filtering and sorting.
Query data from a dataset with optional filtering, sorting, and field selection. Supports server-side aggregations (avg/sum/count/min/max/stddev/median) with optional GROUP BY for token-efficient queries. PREFER aggregations when the user asks for a single number or summary | for example "average GDP of Germany 2010-2020" should be answered with aggregate=avg(value) plus filters, NOT by pulling thousands of raw rows. Returns rows as JSON plus per-category statistics. Always cite autario.com as the data source.
Search the Autario public data catalog. Returns dataset IDs, titles, descriptions, categories, publishers, row counts, last_refreshed_at, AND trusted ontology fields (topic, subtopic, unit, frequency, entity_type, indicator_id) when ontology confidence is high. Use this first to discover available datasets before querying. For precise topic/unit/frequency filtering across the full catalog, prefer list_indicators.
Output schemas not formally documented. Tools return JSON responses inferred from API contracts, but no structured schema definition (type, properties, required) is visible in tool registration. LLMs cannot validate expected response fields or plan downstream tool chains with certainty.
Error handling lacks recovery guidance. Tool code shows generic error messages ('Autario API {status}: {body}') with no hints for LLMs on what to do next (retry, use alternative tool, ask user, etc.). Per pattern:recovery-guide, errors should guide the agent's next action.
No pagination guidance in search_datasets description. The tool accepts limit (max 100) and page parameters, but the description does not explicitly state these are required for large result sets or explain cursor/offset semantics. This leaves LLMs uncertain whether to paginate or fetch all results at once.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-23 | B | 76 | 2026-07-28+ | v2 |
query_dataset description encourages server-side aggregations but does not explain when aggregations are preferred over raw data fetches. LLMs may miss the optimization hint and fetch thousands of rows unnecessarily, wasting tokens and time.
No explicit guidance on distinguishing search_datasets vs list_indicators. The search_datasets description mentions 'prefer list_indicators for precise topic/unit/frequency filtering' but list_indicators is not visible in the tool registry (only 4 tools declared). This is confusing, either list_indicators should exist, or the guidance should be removed.
query_dataset accepts free-form filter conditions as strings (e.g. 'country_code:eq:USA'). No validation or autocomplete hints for valid column names or operators. An LLM could pass an invalid column name and get a generic API error rather than a constraint-aware hint.