Predicting childhood lead exposure at an aggregated level using machine learning
G P Lobo1, B Kalyan1, A J Gadgil1
1Department of Civil and Environmental Engineering, University of California, Berkeley, 94720, United States.
Insights
Machine learning accurately predicts elevated childhood blood lead levels (BLLs) using aggregated data. This approach identifies high-risk areas, enabling targeted interventions for children
Area of Science:
- Environmental Health
- Public Health
- Data Science
Background:
- Childhood lead exposure impacts over 500,000 US children under 6.
- Current universal blood lead screening recommendations are limited.
- Previous predictive models for individual lead exposure had limited success due to data accessibility and geographic variability.
Purpose of the Study:
- To develop and validate a novel machine learning approach for accurately predicting elevated Blood Lead Levels (BLLs) in large groups of children.
- To utilize aggregated, publicly available data for predicting childhood lead exposure.
- To identify geographical hotspots of elevated BLLs for targeted public health interventions.
Main Methods:
- Employed five machine learning models, including Random Forest, to predict childhood lead exposure.
- Utilized socioeconomic, housing, and water quality data aggregated at zip code and city/town levels from New York and Massachusetts.
- Validated the best-performing Random Forest model using New York City data, comparing borough-level predictions with measured BLLs.
Main Results:
- The Random Forest model achieved high performance with 10-fold cross-validation ROC AUC scores of 0.91 (Massachusetts) and 0.85 (New York).
- Model predictions for New York City showed excellent agreement with measured BLLs, predicting an elevated BLL rate of 1.72% compared to the measured 1.73%.
- The study successfully demonstrated the efficacy of using aggregated data for predicting lead exposure hotspots.
Conclusions:
- Machine learning models using aggregated data can accurately predict elevated childhood Blood Lead Levels (BLLs).
- This approach offers a scalable solution for identifying geographical areas with high lead exposure risk.
- The findings support the deployment of targeted public health resources to protect at-risk children.
Abstract:
Childhood lead exposure affects over 500,000 children under 6 years old in the US; however, only 14 states recommend regular universal blood screening. Several studies have reported on the use of predictive models to estimate lead exposure of individual children, albeit with limited success: lead exposure can vary greatly among individuals, individual data is not easily accessible, and models trained in one location do not always perform well in another. We report on a novel approach that uses machine learning to accurately predict elevated Blood Lead Levels (BLLs) in large groups of children, using aggregated data. To that end, we used publicly available zip code and city/town BLL data from the states of New York (n = 1642, excluding New York City) and Massachusetts (n = 352), respectively. Five machine learning models were used to predict childhood lead exposure by using socioeconomic, housing, and water quality predictive features. The best-performing model was a Random Forest, with a 10-fold cross validation ROC AUC score of 0.91 and 0.85 for the Massachusetts and New York datasets, respectively. The model was then tested with New York City data and the results compared to measured BLLs at a borough level. The model yielded predictions in excellent agreement with measured data: at a city level it predicted elevated BLL rates of 1.72% for the children in New York City, which is close to the measured value of 1.73%. Predictive models, such as the one presented here, have the potential to help identify geographical hotspots with significantly large occurrence of elevated lead blood levels in children so that limited resources may be deployed to those who are most at risk.
Related Concept Videos
Steps in Outbreak Investigation
Mechanistic Models: Compartment Models in Individual and Population Analysis
Strategies for Assessing and Addressing Confounding
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
Sampling Plans
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...


