Related Experiment Video
Updated: Aug 8, 2026

Asthma Detection Research Based on Voice Signal Processing and Machine Learning
Published on: July 22, 2025
The crucial role of machine learning models in predicting current childhood asthma: model comparison, calibration,
Aditya Chakraborty1, A K M Raquibul Bashar2
1Department of Epidemiology, Biostatistics and Environmental Health, Joint School of Public Health, Old Dominion University, Norfolk, VA, United States.
Insights
This study developed predictive models for early asthma diagnosis in children. The XGBoost model showed the highest accuracy, while the random forest model offered the best sensitivity for initial screening.
Area of Science:
- Pediatric pulmonology
- Health informatics
- Biostatistics
Background:
- Asthma is a significant chronic childhood illness, particularly challenging to diagnose in young children.
- Predictive models offer potential for early diagnosis, personalized treatments, and understanding disease progression.
- This study leverages national data to build and compare predictive models for asthma.
Purpose of the Study:
- To develop and compare high-performing analytical predictive models for asthma diagnosis.
- To identify key risk factors and influential predictors of asthma in children.
- To improve early detection and management strategies for pediatric asthma.
Main Methods:
- Analysis of the 2011-2020 Behavioral Risk Factor Surveillance System (BRFSS) Asthma Call-Back Survey data (N=9,813).
- Development and comparison of XGBoost, SVM, random forest, LASSO, and GBM models using accuracy, AUC, precision, and recall metrics.
- Evaluation and improvement of model calibration using reliability plots, Platt scaling, and isotonic regression; predictor importance assessed via VIP and SHAP.
Main Results:
- XGBoost demonstrated the highest performance (AUC: 0.95), followed closely by random forest and GBM.
- Random forest achieved the highest sensitivity (0.9786), making it suitable for initial asthma screening.
- Isotonic regression significantly improved model calibration, particularly for the random forest model (ECE reduced from 0.0158 to 0.0086).
- Overnight hospitalization visits and time since last asthma medication were the most influential predictors.
Conclusions:
- The developed analytical methodology effectively identifies behavioral risk factors for asthma.
- Predictive models derived from multidimensional health surveys can aid in diagnosing chronic lung diseases.
- These findings support early asthma diagnosis and proactive management by clinicians.
Background:
Asthma is one of the most prominent chronic diseases in children and one of the most challenging ailments to diagnose in infants and preschoolers in the United States. Predictive models can be instrumental in improving early diagnosis, personalized treatment strategies, and disease progression. By utilizing nationalized data, this study focuses on building and comparing high-performing analytical predictive models based on the relevant risk factors and identifying the most influential predictors.
Methods:
We analyzed cross-sectional BRFSS Asthma Call-Back Survey data (2011-2020; N = 9,813) and randomly split participants into training and testing sets. An XGBoost model (hyperparameters tuned via grid search) was developed and compared with SVM, random forest, LASSO, and GBM using accuracy, AUC, precision, and recall. Calibration was evaluated with reliability plots and improved using Platt scaling and isotonic regression. Predictor contributions were examined using variable-importance (VIP) and Shapley Additive Explanations (SHAP) plot.
Results:
Of the five predictive models, the XGBoost was found to be the best performing model with AUC: 0.95, followed by random forest (AUC: 0.9345), GBM (AUC: 0.9341), SVM (AUC 0.9304), and LASSO (AUC 0.88); however, the random forest model was found to have the highest sensitivity (0.9786), and hence preferred for initial screening of asthma. On the independent test set, calibration (10-bin reliability curves; Brier/ECE/intercept-slope) improved most with isotonic regression, specifically for Random Forest (ECE 0.0158 to 0.0086; intercept -0.174 to -0.010), whereas Platt scaling often worsened calibration, with AUC remaining largely stable across models (AUC ≈ 0.92-0.95). The top two contributing predictors were overnight hospitalization visits and time since the last asthma medication, accounting for 24.62 and 20.92%, respectively, of the asthma status, from the VIP.
Conclusion:
The analytical methodology of model development was found to be instrumental in the discovery of behavioral health-risk knowledge and to visualize the significance of predictive modeling from a multidimensional behavioral health survey. These insights can be instrumental in predicting different types of chronic lung diseases affecting people of all ages and can be useful for clinicians to diagnose asthma at an early stage, allowing for early intervention and proactive management.