MCP server for CDC public health data via Socrata Open Data API (SODA) - Phase 4 Complete: 73 datasets
The CDC MCP server provides a single unified tool with comprehensive parameter coverage across 73 datasets. The tool definition is well-structured with detailed descriptions and an extensive, properly-typed input schema. However, there are significant gaps in output schema documentation, error handling guidance, and parameter descriptions lack actionable detail. The tool naming follows verb_noun convention ('cdc_health_data' is slightly generic but acceptable for a multi-dataset tool). Parameter descriptions exist but many lack specificity about format, range, or constraints. Output structure is undocumented, LLMs cannot plan downstream calls or extract required fields. Error handling provides no recovery guidance. The schema itself is comprehensive (18 parameters with proper enums and type constraints), placing it above the median community server, but the lack of documented output structure and error guidance prevents a higher score.
Unified tool for CDC public health data operations: access disease prevalence, chronic disease indicators, behavioral risk factors, and health surveillance data from CDC's Socrata Open Data API (SODA). Available data sources (73 datasets - Phase 4 Complete: Critical Surveillance): - PLACES: Local disease prevalence data at county, place, census tract, and ZIP code levels - BRFSS: Behavioral Risk Factor Surveillance System for chronic disease risk factors (comprehensive 2011-present) - YRBSS: Youth Risk Behavior Surveillance (substance use, mental health, violence, sexual behaviors) - Respiratory Surveillance: Combined RSV/COVID-19/Flu hospitalization tracking - Vaccination Coverage: Teen (HPV, Tdap, MenACWY), pregnant women, kindergarten immunizations - Birth Statistics: VSRR quarterly birth indicators (rates, preterm, cesarean delivery) - Environmental Health: Air quality tracking (PM2.5, ozone) with health impacts - Tobacco Impact: SAMMEC smoking-attributable mortality, morbidity, economic costs - Oral & Vision Health: NOHSS oral health indicators, BRFSS vision health surveillance - VSRR: Vital Statistics Rapid Release for provisional mortality data - Nutrition/Physical Activity/Obesity: Behavioral and environmental data - Disease-specific datasets: Diabetes, obesity, heart disease, cancer, etc. - NNDSS: National Notifiable Diseases Surveillance (real-time outbreak detection for 50+ diseases) - COVID-19 Vaccination: County-level tracking with equity metrics (SVI, urban/rural) - Drug Overdose: Real-time crisis monitoring with drug-specific tracking (fentanyl, opioids, etc.) Use the method parameter to specify the operation type.
Output schema is completely undocumented. Tool returns results from CDC SODA API but LLMs cannot determine what fields to expect or how to chain results to downstream tools. No documentation of pagination structure, response format, or required fields for follow-up calls.
Parameter descriptions lack actionable constraints. Examples: 'year' described as 'Data release year (e.g., "2023", "2024", or integer for BRFSS)', does not explain valid range, format, or what happens with invalid input. 'state' described only as 'State abbreviation (e.g., "CA", "TX", "NY")', no statement that 2-letter ISO 3166 codes are required, or what to do if an agent passes a full state name.
No error handling guidance. The tool calls external CDC SODA API but provides no documentation of what errors are possible (rate limit, invalid dataset, malformed SoQL), how they are categorized (retryable vs user-fixable vs fatal), or what the LLM should do next. Agents will have no recovery strategy on failure.
| Scored | Grade | Overall | Spec posture | Rubric |
|---|---|---|---|---|
| 2026-09-22 | D | 58 | 2026-07-28+ | v2 |
| 2026-03-09 | F | 36 | - | v1 |
Complex conditional parameter dependencies are poorly documented. Example: 'where_clause' is only valid for method='search_dataset', but this dependency is buried in the description text and not formalized. Agents may pass where_clause to other methods and get confusing errors.
Tool combines 18 methods under one interface. While pragmatic for a multi-dataset CDC tool, this violates the single-responsibility principle and forces the LLM to reason about which method to use based on parameter patterns rather than separate tool names. A tool named 'search_datasets' is clearer than 'cdc_health_data with method=search_dataset'.