Related Experiment Video
Updated: Jan 13, 2026

Author Spotlight: Advancing Alzheimer's Research – Exploring Early Detection and Multi-Omics Approaches
Published on: December 15, 2023
Interpretable Machine Learning Framework for Diabetes Prediction: Integrating SMOTE Balancing with SHAP
Pathamakorn Netayawijit1, Wirapong Chansanam2, Kanda Sorn-In3
1Department of Information Systems, Faculty of Business Administration and Information Technology, Rajamangala University of Technology Isan, Khon Kaen Campus, Khon Kaen 40000, Thailand.
This study enhances machine learning for diabetes prediction by combining Synthetic Minority Oversampling Technique (SMOTE) with SHapley Additive exPlanations (SHAP). The Random Forest-SMOTE model achieved high accuracy and identified key predictors like glucose and BMI.
Area of Science:
- Machine Learning in Healthcare
- Diabetes Prediction Models
- AI for Clinical Decision Support
Background:
- Class imbalance and limited interpretability hinder clinical adoption of AI for diabetes prediction.
- Existing models often show poor sensitivity to high-risk cases and lack clinical trust.
- This study integrates SMOTE resampling with SHAP explainability to improve performance and transparency.
Purpose of the Study:
- To develop and validate an interpretable machine learning framework for diabetes prediction.
- To address class imbalance using advanced resampling techniques (SMOTE).
- To provide clinically meaningful explanations via SHAP for enhanced decision support.
Main Methods:
- A seven-stage pipeline combining Synthetic Minority Oversampling Technique (SMOTE) and SHapley Additive exPlanations (SHAP).
- Comparative evaluation of five algorithms (Random Forest, Gradient Boosting, SVM, Logistic Regression, XGBoost) on 1500 patient records.
- 5-fold stratified cross-validation with SMOTE applied within training folds to prevent data leakage.
Main Results:
- The Random Forest-SMOTE model achieved 96.91% accuracy, 0.998 AUC, 99.5% sensitivity, and 97.3% specificity.
- SHAP analysis identified glucose (SHAP value: 2.34) and BMI (SHAP value: 1.87) as primary predictors.
- Feature interaction analysis revealed synergistic effects between glucose and BMI.
Conclusions:
- The proposed framework shows promise in bridging algorithmic performance and clinical applicability for diabetes prediction.
- The Random Forest-SMOTE model demonstrated high cross-validated performance on a publicly available dataset.
- Further external, prospective validation in real-world cohorts is required before clinical deployment.
Related Concept Videos
Diabetes Mellitus: Overview and Type I Subtype
Type 1 diabetes is an autoimmune disease in which the immune system mistakenly attacks and destroys the insulin-producing beta cells in the pancreas. As a result, the body is unable to produce sufficient insulin, and individuals with...
Diabetes: Management and Pharmacotherapy
Insulin remains the cornerstone of treatment for most patients with type 1 and many...
Diabetes Mellitus: Type 2 and Gestational
Pathophysiology of Diabetes
Type 1 diabetes is characterized by autoimmune-mediated destruction of pancreatic β cells, with environmental factors potentially triggering this process in genetically susceptible individuals. Despite many not having a family history, certain genes increase susceptibility,...