Related Experiment Videos
Development and validation of an explainable machine learning model for differentiating diabetic nephropathy from
Yonglin Zhang1, Siyu Feng2, Yukun Xue3
1Department of Pharmacy, Affiliated Hospital of North Sichuan Medical College, Nanchong, Sichuan, China.
Background:
Diabetic nephropathy (DN) and diabetic retinopathy (DR) are common microvascular complications of type 2 diabetes mellitus (T2DM) and may require different diagnostic and management pathways. This study aimed to develop and validate an interpretable machine learning model based on routine laboratory data to differentiate prevalent DN from prevalent DR among hospitalized patients with type 2 diabetes.
Methods:
Data were collected from a large tertiary hospital in China and split into a training/internal validation cohort (DN: 2,309 cases; DR: 855 cases) and an independent held-out validation cohort (DN: 578 cases; DR: 214 cases). A total of 47 routinely available laboratory and demographic variables were extracted from electronic health records (EHRs). Seven machine learning algorithms were developed and compared, with recursive feature elimination (RFE) employed to identify the most informative subset of features and enhance model performance and interpretability. Model discrimination was assessed using the area under the receiver operating characteristic curve (AUC) and the area under the precision-recall curve (AP), while SHAP values were used to interpret feature importance and explain individual-level predictions.
Results:
The extreme gradient boosting (XGBoost) classifier demonstrated the highest predictive performance among the seven machine learning algorithms evaluated. After selecting the top five features based on importance rankings, an explainable XGBoost model was constructed. This final model achieved strong apparent discrimination in both the training/internal validation cohort (AUC = 0.991, 95% CI: 0.989-0.994; AP = 0.979, 95% CI: 0.973-0.984) and the held-out validation cohort (AUC = 0.997, 95% CI: 0.996-0.999; AP = 0.993, 95% CI: 0.988-0.997). SHAP analysis further identified α-hydroxybutyrate dehydrogenase, creatine kinase-MB, creatinine, urinary α1-microglobulin, and N-acetyl-β-D-glucosaminidase as the most influential features contributing to complication risk prediction.
Conclusions:
An explainable machine learning model for predicting complications in patients with T2DM demonstrated high feasibility and effectiveness, indicating strong potential to support clinical management and improve patient outcomes. By incorporating SHAP analyses, the model addresses key concerns regarding transparency and clinical decision-making. These findings highlight the model's potential for real-world clinical implementation.
Related Concept Videos
Diabetic Nephropathy
Diabetic Retinopathy
Diabetes Mellitus: Type 2 and Gestational
Type II Diabetes Mellitus III: Clinical Manifestations and Diagnosis