MCP server wrapping the CDC WONDER Data Query Web Service for US public health statistics: mortality, natality, cancer, and more.
CDC WONDER server has 14 well-organized tools with mostly solid descriptions and input schemas. Naming is verb-forward and clear. However, there are systematic gaps: output schemas are almost never documented, error handling lacks recovery guidance, parameter descriptions are sometimes vague, and some tools accept overly generic dict inputs without clear structure. The server follows good naming conventions (verb_noun pattern like list_databases, query_wonder, calculate_*) and most tools have helpful docstrings explaining what they do and when to use them. However, critical improvements are needed in output documentation, error classification, and input validation clarity.
Build a standard US life table (remaining life expectancy by age) from CDC WONDER mortality data and population estimates. Workflow: 1. Query WONDER for age-specific death rates and populations 2. Provide age_groups (e.g. ["0-4", "5-14", ...]), death_counts, populations 3. Tool computes qx (probability of death), lx (survivors), Lx, Tx, ex (remaining life expectancy) 4. Returns full life table for use in YLL or other calculations
Compute Annual Percent Change (APC) for a rate time series. Fits log(rate) ~ year via OLS. Returns APC, 95% CI, p-value, and trend label. This is the single-segment APC (not joinpoint regression). Typical workflow: 1. Call query_wonder with group_by=["D76.V1-level1"] and measures=["D76.M3"] or ["D76.M4"] 2. Extract the year and rate columns from the result rows 3. Pass here as years and rates
Estimate excess deaths by comparing observed mortality to a counterfactual baseline. Baseline can be a historical period, a reference population, or a model fit. Returns the difference and its 95% CI under Poisson assumptions. Workflow: 1. Query observed counts for a recent period (e.g. 2020–2021) and a baseline period (e.g. 2015–2019) 2. Compute observed total (recent) and expected total (baseline or projected) 3. Pass to calculate_excess_deaths 4. Results: excess_deaths and 95% CI
Decompose a rate difference into components attributable to age-stratum-specific rate differences vs. age composition differences (Kitagawa decomposition). Useful for determining whether a mortality difference between two populations is driven by: - Different rates within each age group ("rate component") - Different age distributions ("composition component") Workflow: 1. Query stratum-specific rates (e.g. by age) for two populations 2. Gather observed counts and populations per stratum for each group 3. Call calculate_kitagawa 4. Outputs: overall RD, rate component, composition component, and % attributable to each
Output schemas almost entirely undocumented. Tools return complex dicts (e.g., query_wonder returns QueryResult.model_dump(), calculate_smr returns CI bounds and SMR value) but LLMs have no visible schema describing the field names, types, and structure. This forces LLMs to guess what fields to expect and breaks downstream tool chaining.
Generic dict parameters lack internal structure documentation. Tools like calculate_rate_ratio, calculate_rate_difference, calculate_smr, calculate_kitagawa, and calculate_yll accept dicts with keys like 'count', 'population', 'rate', 'rate_per', 'rate_se', 'label' but the parameter description only says 'Dict with keys: ...' without specifying types (int vs float), whether fields are required or optional, or what happens if incompatible keys are provided. LLMs cannot validate inputs before calling.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | C | 69 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 0 | - | v1 |
Compute absolute rate difference (group_1 - group_2) with CI. CI method selected based on available inputs: 1. count + population → Poisson variance per group 2. rate + rate_se → standard error propagation 3. rate only → difference computed, CI not available
Compute a rate ratio (group_1 / group_2) with a confidence interval. This is a fully deterministic calculation — no API call is made. Uses the Poisson exact method per group and delta method for the ratio CI. Each group must provide EITHER: - count (int) + population (int): exact CI computed - rate (float) + rate_per (int, default 100000): CI not computable
Compute standardized mortality ratio (SMR) for a small population against the national standard. For a given cause and age-adjusted population reference, SMR = observed / expected. Useful for comparing a sub-population (e.g. occupational cohort) to national rates. Workflow: 1. Query national rates by age stratum (group_by=["D76.V5"] for age groups) 2. For your cohort, provide observed deaths by age stratum 3. Pass expected_rate_per_stratum (national rates) and observed_deaths_per_stratum 4. SMR computed with Poisson-based CI (exact when possible)
Compute years of life lost (YLL) attributable to a cause of death. YLL = Σ (deaths_in_age_group × remaining_life_expectancy_at_age). Uses a standard US life table (provided or user-supplied). Returns total YLL, YLL per death, and YLL rate per 100,000 population.
Generate a self-contained Python replication script from a QueryResult. The script embeds the exact XML and requires only requests + beautifulsoup4. It can be run in any Python environment without this MCP server.
Generate a self-contained R replication script from a QueryResult. The script embeds the exact XML and requires only httr + xml2. It does NOT depend on the wonderapi package or this MCP server.
Return the group-by variables, measures, and filter variables for a database. Use this to discover valid parameter codes before calling query_wonder.
Look up ICD-10-CM diagnosis codes by keyword, chapter, or sub-chapter. Searches a local ICD-10 registry (not an API call). Useful for constructing WONDER filters on cause of death. Examples: - get_icd10_codes("heart disease") → ICD-10 codes for cardiovascular conditions - get_icd10_codes(chapter="Chapter IX") → All circulatory system codes - get_icd10_codes(code_range="I20-I25") → Ischemic heart disease codes
Return the registry of known CDC WONDER databases. No API call is made — this is a local lookup. Includes database IDs, labels, year ranges, and brief descriptions.
Execute a CDC WONDER query and return a structured result. The result includes the data table, caveats, the exact XML sent, and a timestamp — everything needed to reproduce the query. IMPORTANT: The API enforces a ~2-minute rate limit between requests. This tool will automatically wait if called too soon after a previous query. Geographic limitation: Only national-level data is available via the API. Sub-national data (state, county) requires the WONDER web interface.
Error handling does not guide recovery. query_wonder returns {"error": str(e), "raw_response_preview": ...} but does not indicate whether the error is retryable, user-fixable (e.g. invalid database_id, filter syntax), or fatal. No recovery guidance like 'Use list_databases() to find valid IDs' is provided in error returns.
No input validation or constraint enforcement visible in code. Parameters like group_by (max 5 items per description) and measures are lists with no length checks; filters and options are dicts with no validation. LLMs can pass invalid values (6 group_by items, unknown filter keys) and fail silently or with cryptic API errors.
Parameter descriptions mention expected values and formats (e.g., 'Example: ["D76.V1-level1", "D76.V8"]') but do not formalize them as enums or patterns. LLMs may treat examples as the only valid values or hallucinate variations.
Vague parameter descriptions in statistics tools. 'Dict with keys: count, population, rate, rate_per (int), rate_se (float), label (str)' in calculate_rate_ratio and similar does not clarify: (1) Are all keys required? (2) What if only 'rate' is provided, does CI calculation skip? (3) What type is 'label'? Descriptions should state the dependency logic explicitly.