MCP server providing tools to interact with IPUMS microdata and NHGIS extract APIs, including metadata browsing, extract creation, code generation, and file downloads
The server demonstrates solid definition quality with 22 well-named tools covering two distinct data domains (IPUMS microdata and NHGIS). Tool names follow verb_noun conventions (list_, get_, create_, search_, extract_to_code) and clearly indicate intent. All tools have descriptions of adequate length (typically 100-200 chars). Input schemas are visible and properly typed with enums for collection/dataset parameters and string descriptions for most fields. However, there are systematic gaps: (1) output schemas are not documented, callers cannot see what fields are returned; (2) parameter descriptions are often generic or lack validation constraints; (3) error handling guidance is absent from descriptions; (4) some complex object parameters (dataStructure, extract, datasets) lack internal schema documentation. These gaps are noticeable but not fatal, the server is functional and usable by agents, just not optimized for LLM reasoning about return values and error recovery.
Submit a new IPUMS microdata extract request. Specify the collection, samples, variables, and output format. Returns the new extract number and initial status. Once submitted, use microdata_extract_to_code to generate reproducible R or Python code for the extract.
Download completed IPUMS microdata extract files to a local directory. Verifies SHA-256 checksums. Returns local file paths and verification status.
Generate reproducible R or Python code for a microdata extract. Takes an extract specification and outputs ready-to-run code that recreates the extract, submits it, waits for completion, downloads files, and loads data into the analysis environment.
Get the status and details of a specific IPUMS microdata extract by its number. Includes download links when the extract is complete.
List recent extracts for an IPUMS microdata collection (e.g. usa, cps, ipumsi). Returns extract numbers, status, and basic metadata.
Output schemas are not documented. Tools like microdata_list_extracts, microdata_get_extract, nhgis_get_dataset do not describe what fields are returned. LLMs cannot reason about the response structure or chain tools (e.g., what fields are available to pass to downstream tools). This violates pattern:tool which requires documented return types.
Complex object parameters lack internal schema documentation. microdata_create_extract and nhgis_create_extract accept 'dataStructure', 'extract', 'datasets' as object types without documenting their internal structure. LLMs cannot know what fields to pass inside these objects. Schema should document nested properties with types and descriptions.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 67 | 2026-07-28+ | v2 |
| 2026-03-09 | C | 62 | - | v1 |
List available samples for an IPUMS microdata collection (e.g. all ACS, CPS, or Census years). Use this to find the correct sample ID before creating an extract.
Search for microdata variables by name or keyword. For USA microdata, searches the local USA_VARIABLES dataset (827+ harmonized variables). Returns matching variable names, labels, types (H=household, P=person), thematic groups, and available sample years.
Poll an IPUMS microdata extract until it reaches 'completed' or 'failed' status. Returns the final extract details. Useful for small/medium extracts; for large extracts consider checking back manually with microdata_get_extract.
Submit a new NHGIS extract request. Specify datasets (with data tables and geographic levels), time series tables, shapefiles, and output format. Returns the new extract number and initial status. Once submitted, use nhgis_extract_to_code to generate reproducible R or Python code for the extract.
Generate reproducible R or Python code for an NHGIS extract. Takes an extract specification and outputs ready-to-run code that recreates the extract, submits it, waits for completion, downloads files, and loads data into the analysis environment.
Get detailed information about a specific NHGIS data table within a dataset, including all variable descriptions.
Get full details for a specific NHGIS dataset including available data tables, geographic levels, breakdown values, and years.
Get the status and details of a specific NHGIS extract by its number. Includes download links when the extract is complete.
Get full details for a specific NHGIS time series table including available years, geographic levels, and variable descriptions.
List all available NHGIS data tables across all datasets. Returns paginated list of table identifiers, descriptions, and universe.
Browse available NHGIS datasets. Returns paginated list of dataset identifiers, names, years, and geographic levels available.
List recent NHGIS extract requests. Returns extract numbers, status, and submission timestamps.
List all available NHGIS shapefiles (boundary files). Returns identifiers for census geographies that can be included in extracts.
List all available NHGIS time series tables. These span multiple census years with consistent geographic definitions.
Search NHGIS data tables by keyword. Fetches all available data tables and returns those whose identifier, description, or universe match the keyword. Use this to find tables for a topic before creating an extract.
Search NHGIS datasets by keyword. Fetches all available datasets and returns those whose name, group (e.g. '2020 Census'), or description match the keyword. Use this to find which datasets cover a topic or census year before drilling into specific data tables.
Search NHGIS time series tables by keyword. Fetches all available time series tables and returns those whose name, description, or universe match the keyword. Use this to find time-series data before creating an extract.
No error handling guidance in tool descriptions. Tools like microdata_create_extract and nhgis_create_extract do not describe what errors can occur (e.g., invalid collection, missing required variables, API rate limits) or how to recover. LLMs cannot self-correct or retry intelligently when calls fail.
Parameter descriptions lack validation constraints. For example, 'maxLength' is declared for search keywords (200 chars) but many numeric parameters (pageNumber, pageSize, maxWaitSeconds, pollIntervalSeconds) lack min/max bounds in descriptions. LLMs may pass invalid values like pageNumber=0 or pageSize=10000.
Tool descriptions do not indicate which tools read-only vs which write/modify state. While the provided metadata marks some as READ_ONLY or WRITE, the human-facing descriptions do not explicitly state 'This operation is permanent and cannot be undone' or 'This is read-only, safe to retry.' Agents need this for planning.