Related Experiment Video
Updated: Jun 3, 2025

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
Leveraging Shapley Additive Explanations for Feature Selection in Ensemble Models for Diabetes Prediction
Prasant Kumar Mohanty1, Sharmila Anand John Francis2, Rabindra Kumar Barik3
1Department of Computer Science and Engineering, National Institute of Technology, Meghalaya 793003, India.
This study uses Shapley Additive explanations (SHAPs) with machine learning to identify key diabetes risk factors. Focusing on the top three features significantly improved prediction accuracy and efficiency for clinical applications.
Area of Science:
- Medical Informatics
- Computational Biology
- Public Health
Background:
- Diabetes poses a global health challenge, exacerbated in India by lifestyle changes linked to urbanization.
- Effective diabetes management requires advanced prevention strategies and technological integration.
Purpose of the Study:
- To enhance the accuracy and efficiency of diabetes prediction models.
- To identify and validate the most influential features for diabetes risk using SHAP values.
- To assess the impact of feature selection on model performance.
Main Methods:
- Integration of Shapley Additive explanations (SHAPs) with ensemble machine learning models.
- Evaluation of model performance using three feature sets: all features, top three influential features, and features excluding the top three.
- Comparative analysis of ten different machine learning models.
Main Results:
- Models utilizing the top three most influential features demonstrated superior performance.
- The ensemble model achieved better performance across most metrics when focusing on the top features.
- Excluding the top three features resulted in a significant decrease in predictive performance.
Conclusions:
- Targeted feature selection using SHAP values is effective for improving diabetes prediction models.
- This approach enhances model efficiency and robustness for clinical applications.
- Identifying key predictive features is crucial for accurate and resource-efficient diabetes management.
Related Concept Videos
Sensitivity, Specificity, and Predicted Value
Sensitivity is the...
Diabetes Mellitus: Type 2 and Gestational
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Diabetes Mellitus: Overview and Type I Subtype
Type 1 diabetes is an autoimmune disease in which the immune system mistakenly attacks and destroys the insulin-producing beta cells in the pancreas. As a result, the body is unable to produce sufficient insulin, and individuals with...
Cancer Survival Analysis
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.

