Prediction of measles cases in US counties: A machine learning approach

Stephanie A Kujawski1, Boshu Ru1, Nelson Lee Afanador1

  • 1Merck & Co., Inc. Rahway, NJ, USA.

Vaccine
|September 7, 2024
PubMed
Abstract

Insights

A machine learning model accurately predicted measles outbreaks in most US counties. This tool can help health agencies prevent and contain future measles cases effectively.

Area of Science:

  • Epidemiology
  • Public Health
  • Machine Learning in Healthcare

Background:

  • Measles, declared eliminated in the US in 2000, has seen a rise in outbreaks.
  • Predicting measles case locations is crucial for effective prevention and containment strategies.

Purpose of the Study:

  • To develop and validate a machine learning model for predicting county-level measles risk in the United States.
  • To aid public health agencies in identifying high-risk areas for measles prevention.

Main Methods:

  • A machine learning model was developed using 17 predictor variables.
  • The model was trained on US county-level measles case data from 2014 and 2018.
  • Model performance was evaluated by comparing predictions with actual 2019 measles case data.

Main Results:

  • The model achieved 95% specificity and 72% sensitivity in predicting counties with and without measles cases in 2019.
  • It accounted for 94% of all reported measles cases in 2019.
  • The model correctly identified 73% of the top 30 highest-risk counties, capturing 72% of all cases.

Conclusions:

  • The developed machine learning model accurately predicts a majority of US counties at high risk for measles.
  • This predictive framework can significantly support state and national health agencies in their measles prevention and containment efforts.

Related Concept Videos

Steps in Outbreak Investigation01:18

Steps in Outbreak Investigation

In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
114
Prediction Intervals01:03

Prediction Intervals

The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y. 
2.2K
Statistical Methods for Analyzing Epidemiological Data01:25

Statistical Methods for Analyzing Epidemiological Data

Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
330
Residuals and Least-Squares Property01:11

Residuals and Least-Squares Property

The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.3K
Classification of Leukocytes01:30

Classification of Leukocytes

Leukocytes are classified into two groups based on the presence or absence of cytoplasmic granules. Granular leukocytes, which contain granules, belong to the myeloid lineage and are divided into three subtypes: neutrophils, eosinophils, and basophils. These cells are roughly spherical and characterized by the granules in their cytoplasm.
Neutrophils are the most abundant type of granular leukocytes, comprising 50-70% of all leukocytes. They feature small, evenly distributed granules and a...
1.8K
End Point Prediction: Gran Plot01:07

End Point Prediction: Gran Plot

A Gran plot is used to predict the equivalence volume or endpoint of a potentiometric or acid-base titration without reaching the endpoint. Typically, titration data is collected as a function of the titrant's volume up to a point less than the equivalence volume and then transformed into a linear format. The straight line is extended to the x-axis, indicating the necessary titrant volume to achieve the equivalence point.
For potentiometric titration, the Gran plot is created by plotting...
300