Comparative Evaluation of Explainable Machine Learning Versus Linear Regression for Predicting County-Level Lung
Soheil Hashtarkhani1, Brianna M White1, Benyamin Hoseini2
1Department of Pediatrics, Center for Biomedical Informatics, College of Medicine, University of Tennessee Health Science Center, Memphis, TN.
Machine learning accurately predicts lung cancer mortality. Smoking rates, home values, and Hispanic population percentage are key factors, with disparities concentrated in the mid-eastern US.
Area of Science:
- Public Health
- Epidemiology
- Machine Learning
Background:
- Lung cancer (LC) is a major cause of cancer mortality in the U.S.
- Accurate LC mortality prediction is vital for interventions and addressing health disparities.
- Explainable machine learning offers potential for improved prediction and insight over traditional models.
Purpose of the Study:
- To predict county-level lung cancer mortality rates across the U.S.
- To compare the performance of random forest (RF), gradient boosting regression (GBR), and linear regression (LR) models.
- To identify key factors influencing LC mortality and analyze geographic disparities.
Main Methods:
- Applied RF, GBR, and LR models to predict U.S. county-level LC mortality.
- Evaluated model performance using R-squared and root mean squared error (RMSE).
- Utilized Shapley Additive Explanations (SHAP) for variable importance and Getis-Ord (Gi*) for spatial hotspot analysis.
Main Results:
- The RF model achieved the highest predictive accuracy (R-squared: 41.9%, RMSE: 12.8).
- Key predictors included smoking rate, median home value, and Hispanic population percentage.
- Significant LC mortality hotspots were identified in the mid-eastern U.S. counties.
Conclusions:
- The RF model provides superior prediction of LC mortality rates.
- Smoking prevalence, housing values, and Hispanic population percentage are critical factors.
- Findings support targeted interventions and health disparity reduction strategies in high-risk areas.
More Related Videos
07:13Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model
Published on: April 18, 2025
04:09Predicting Treatment Response to Image-Guided Therapies Using Machine Learning: An Example for Trans-Arterial Treatment of Hepatocellular Carcinoma
Published on: October 10, 2018
Related Concept Videos
Statistical Methods for Analyzing Epidemiological Data
Residuals and Least-Squares Property
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
Comparing the Survival Analysis of Two or More Groups
Cancer Survival Analysis
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Kaplan-Meier Approach
