Model Context Protocol (MCP) server that gives LLM agents and chatbots direct access to the World Bank's Data360 Platform. Search, validate, and retrieve development indicators—from GDP and poverty to gender equality and climate—with structured metadata and time-series data, without hallucinating values.
Server exhibits mixed quality across 10 tools. Strengths: all tools have descriptions (well above 20 chars), input schemas are visible with type declarations, and naming is consistent with verb_noun convention (data360_search_*, data360_get_*, data360_create_*, etc.). Weaknesses: (1) Parameters frequently lack detailed constraints (enums, ranges, format specs), descriptions say 'Optional' or 'Max X' but omit validation rules LLMs can enforce; (2) Output schemas are NOT documented, no tool definition shows what fields are returned, making it impossible for LLMs to plan downstream calls or extract data reliably; (3) Error handling is not visible in tool definitions, no recovery guidance, no categorization of retryable vs fatal errors; (4) Some tools combine multiple concerns (e.g. data360_get_data does retrieval + optional disaggregation filtering + pagination, which could be decomposed); (5) Tool descriptions are action-focused but lack explicit 'when to use' guidance relative to similar tools (e.g., when to call data360_search_indicators vs data360_search_datasets). Average tool definition score: 68 (naming 85/100, description 72/100, schema 52/100 due to missing output specs).
Aggregate indicator data across dimensions or countries using configurable methods.
Create a Vega-Lite chart visualization from Data360 indicator data.
Expand a country group (e.g., regional or income group) to get individual country ISO codes.
Find valid codes for dimensions (e.g., country names to ISO codes, or dimension values).
Retrieve indicator observations from the Data360 API. Use when you need actual numeric values (OBS_VALUE) for specific countries and years. Ensure the database ID and the indicator ID are already in context before using this tool. Do not guess or hallucinate these IDs. Call `data360_get_disaggregation` first to find available years and breakdowns for the `disaggregation_filters`.
Get available disaggregation dimensions (breakdowns) and their values for an indicator.
Output schemas are completely undocumented. Tool definitions show input parameters but no return-value structure. LLMs cannot infer what fields are returned (e.g., does data360_search_indicators return 'indicator_id' or 'id'? Is there a 'database_id' field?). This forces agents to guess field names for downstream chaining and prevents confident data extraction.
Parameter descriptions lack actionable validation constraints. Many params are described as 'Optional' or 'Max X records' without specifying enums, ranges, formats, or character restrictions. Example: 'limit' is described as 'Max datasets to return (default 10)' but lacks min/max bounds (is 1 - 100 valid? 1 - 1000?). Descriptions should state 'integer, range 1 - 100' not just 'integer'.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 63 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 52 | - | v1 |
Get metadata and disaggregation options for a Data360 indicator. Use when you need detailed methodology, source notes, or limitations for an indicator. Ensure the database ID and the indicator ID are already in context (e.g., from `data360_search_indicators`) before using this tool. Do not guess or hallucinate these IDs.
Search for Data360 datasets matching a query. Use when the user asks for dataset details, catalogs, or source databases (e.g. 'Findex', 'WDI').
Search for Data360 indicators with enriched metadata for selection. Use when the user asks for data on a development topic (e.g. GDP, poverty, education). Default to using the single `query` parameter for any single topic/indicator search. Use the `queries` or `query_groups` parameters ONLY when the request involves multiple topics or scopes (2 or more).
Validate indicator data quality and check for missing values, outliers, or data coverage issues.
No error handling or recovery guidance in tool definitions. Tools do not document what errors can occur, whether they are retryable, or what the LLM should do next. E.g., if database_id is invalid, does the tool return 'not_found' or 'invalid_parameter'? Should the LLM retry, ask the user, or call a discovery tool?
Parameters that depend on prior discovery calls lack explicit dependency hints. For example, data360_get_data requires database_id and indicator_id but the description does not state 'first call data360_search_indicators to find these IDs'. Similarly, disaggregation_filters references 'Call data360_get_disaggregation first' in passing but this is not prominently featured. LLMs often skip discovery steps, causing invalid calls.
Tool descriptions lack clear 'when to use' guidance relative to similar tools. For example, data360_search_indicators and data360_search_datasets both search but the descriptions do not explicitly state: 'Use search_indicators when you need specific indicators (e.g., GDP); use search_datasets when you need dataset catalogs or sources.' Ambiguous boundaries cause LLMs to call both tools wastefully.
Some tools combine multiple concerns or workflows. data360_get_data accepts optional disaggregation_filters but the description does not make clear that calling data360_get_disaggregation first is prerequisite to knowing valid filter values. This couples the tools tightly and could be clarified with a dedicated 'disaggregation lookup' step. Consider whether data360_aggregate is distinct enough from data360_get_data with aggregation applied.
Return value cardinality not specified. Tools like data360_expand_country_group and data360_find_codelist_value do not state whether results are paginated, how many items are returned by default, or whether pagination is supported. If a search returns 500 results, does the tool cap at a default limit or return all? Agents need this to plan response handling.
Parameter 'queries' (array) in data360_search_indicators can accept 'at least 2 non-empty search strings' but the description does not explain what happens if you pass 1 or 3+ items. Is the minimum enforced? Maximum? Does it error or silently ignore? Type is 'array|null' but no item schema is visible (strings? objects?).
query_groups parameter in data360_search_indicators has structure [{'queries': [...], 'country': ...}] but the schema does not show the inner structure, are queries always strings? Is country a string or array? What fields are required vs optional? The description is vague about the expected object shape.