Model Context Protocol (MCP) server for interacting with data.gouv.fr datasets and resources via LLM chatbots
The data.gouv.fr MCP server demonstrates solid definition quality with consistent naming patterns, clear descriptions for all tools, and well-structured input schemas. All 12 tools follow verb_noun naming conventions (search_, get_, list_) and every tool has a non-trivial description. Input parameters are typed and described. However, the server lacks explicit output schema documentation in the visible code, missing per-tool error guidance, and does not evidence per-parameter constraint documentation (enum values, ranges, format specs). Tool annotations are correctly applied (readOnlyHint, idempotentHint, openWorldHint across all tools), demonstrating awareness of tool metadata patterns. The parameter quality is good (avg 5 params per tool) with sensible defaults (page_size defaults to 20/100, sort_order defaults to 'desc'). No security issues detected (no credentials in parameters). The main gap is lack of visible output schema definitions and error recovery guidance in tool descriptions.
Get metadata for a third-party API including id, title, description, organization, documentation URL, and tags.
Get metadata for a single dataset including dataset ID, title, short and long descriptions.
Fetch metrics for a given dataset (views, downloads, reuses, followers) with specified time granularity.
Fetch metrics for a given organization with specified time granularity.
Fetch paginated rows from a tabular resource via the Tabular API with optional filtering and sorting.
Get metadata for a single resource including resource ID, title, description, and associated dataset ID.
Fetch metrics for a given resource (views, downloads, reuses, followers) with specified time granularity.
Output schemas not documented in tool definitions. While input schemas are present and typed, there is no visible documentation of what fields each tool returns (e.g., search_datasets should document that it returns dataset IDs, titles, descriptions, and metadata structure). LLMs cannot plan downstream calls or extract data without knowing output field names and types.
Parameter constraints are not formalized in descriptions. For example, sort_last_update_range accepts only 'last_30_days', 'last_12_months', 'last_3_years' but this is not declared as an enum constraint in the visible schema. Similarly, sort_order accepts 'asc' or 'desc' but lacks enum enforcement. This invites hallucinated values from LLMs.
Inferred effective spec: 2026-07-28+.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | B | 70 | 2026-07-28+ | v2 |
| 2026-03-09 | D | 57 | - | v1 |
Get the schema (column names and types) for a tabular resource via the Tabular API profile endpoint.
Fetch metrics for a given reuse with specified time granularity.
Get all resources for a given dataset with their IDs and titles.
Search third-party APIs cataloged on data.gouv.fr via GET /2/dataservices/search/
Search for datasets on data.gouv.fr. Searches in title, description, tags. Accepts pagination, sorting, and date filtering.
Error handling guidance missing from tool descriptions. None of the 12 tools provide recovery paths in their descriptions (e.g., 'If resource not found, try search_datasets first' or 'On 4xx error, verify the ID is valid'). This forces LLMs to guess the next step when calls fail.
Pagination limits not consistently documented in descriptions. search_datasets and search_dataservices mention max page_size=100, but get_resource_data says 'default: 100' without stating a maximum. Unbounded pagination can cause token exhaustion or API timeouts.