Predicting vaccine hesitancy from area-level indicators: A machine learning approach.
Vincenzo Carrieri1,2,3, Raffele Lagravinese4, Giuliano Resce5
1Department of Law, Economics and Sociology, Magna Graecia University, Catanzaro, Italy.
Health Economics
|September 15, 2021
Summary
Machine learning models can predict communities at high risk of vaccine hesitancy (VH). This approach improves accuracy by 24%, identifying waste recycling and employment rates as key predictors for targeted public health interventions.
Area of Science:
- Public Health
- Epidemiology
- Health Policy
Background:
- Vaccine hesitancy (VH) poses a significant challenge to public health initiatives, potentially undermining mass immunization campaigns like those for COVID-19.
- Predicting and understanding the drivers of VH at a community level is crucial for developing effective public health strategies.
- Existing methods for identifying at-risk communities often lack the predictive power needed for timely intervention.
Purpose of the Study:
- To develop and evaluate machine learning models for predicting areas with high vaccine hesitancy.
- To identify key area-level indicators that are most strongly associated with vaccine hesitancy.
- To provide policymakers with a tool to target public health interventions and awareness campaigns effectively.
Main Methods:
- Utilized machine learning algorithms to analyze data from child immunization campaigns for seven non-mandatory vaccines across 6062 Italian municipalities in 2016.
- Compared the predictive performance of various machine learning models using the area under the receiver operating characteristics curve (AUC).
- Identified significant area-level predictors of vaccine hesitancy from a range of socioeconomic and environmental indicators.
Main Results:
- The Random Forest algorithm demonstrated the highest predictive accuracy for identifying areas at high risk of VH.
- The model achieved a 24% improvement in accuracy compared to the baseline unpredictable level.
- The proportion of waste recycling and the employment rate were identified as the most potent predictors of high vaccine hesitancy.
Conclusions:
- Machine learning offers a powerful approach to predict community-level vaccine hesitancy using readily available data.
- Area-level indicators, such as waste recycling and employment rates, can serve as valuable proxies for identifying populations at risk of VH.
- These findings can inform the strategic deployment of public health resources and targeted pro-vaccine awareness campaigns to mitigate vaccine hesitancy.
Related Concept Videos
Steps in Outbreak Investigation
259
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
259
Prediction Intervals
2.5K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.5K
Residuals and Least-Squares Property
8.1K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
8.1K
Bias in Epidemiological Studies
817
Biases can arise at various stages of research, from study design and data collection to analysis and interpretation. Recognizing and addressing these biases is essential to ensure the validity and reliability of epidemiological findings.Broadly speaking, biases in epidemiology fall into three main categories: selection bias, information bias, and confounding. A more detailed description of possible biases is:
817
Statistical Methods for Analyzing Epidemiological Data
625
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
625
Sensitivity, Specificity, and Predicted Value
816
In healthcare diagnostics, laboratory tests play a crucial role in identifying and diagnosing a wide range of medical conditions. However, interpreting test results is not always straightforward. An abnormal test result does not always confirm the presence of a disease, just as a normal result does not guarantee its absence. To assess the reliability of these diagnostic tools, healthcare practitioners rely on two key statistical indicators: sensitivity and specificity.
Sensitivity is the...
Sensitivity is the...
816


