Related Experiment Videos
Machine learning versus traditional regression models for predicting diabetic retinopathy screening adherence among
Yawen Wang1, Lanting Xia1, Guiling Geng2
1School of Nursing and Rehabilitation, Nantong University, Nantong, China.
Background:
Diabetic retinopathy (DR) is a leading cause of preventable vision loss among patients with diabetes, yet screening adherence remains suboptimal. Existing studies have mainly focused on clinical or sociodemographic determinants, with limited evidence integrating psychological mechanisms. The application of Protection Motivation Theory (PMT)-based constructs combined with machine learning for predicting DR screening adherence remains underexplored in community populations.
Objective:
This study compared machine learning models (decision tree and random forest) with traditional logistic regression for predicting DR screening adherence among community-dwelling older adults with diabetes, incorporating a validated PMT-based questionnaire as a key psychological predictor.
Methods:
A cluster random sampling design recruited 1,021 older adults with diabetes from four community health centers in Nantong, China (March-October 2025). Data included sociodemographic characteristics, clinical indicators, health behaviors, and PMT-based constructs. Participants were classified as good (n = 159) or poor adherence (n = 862). SMOTE and class weighting addressed class imbalance. Models were evaluated using 10-fold cross-validation. Logistic regression, decision tree, and random forest models were built using identical predictors. Decision curve analysis assessed clinical utility.
Results:
The random forest model achieved a marginally higher AUC (0.774, 95% CI: 0.678-0.869) compared with logistic regression (AUC = 0.751, 95% CI: 0.678-0.824) and decision tree (AUC = 0.709, 95% CI: 0.596-0.821); however, DeLong's tests indicated no statistically significant differences (all p > 0.05). The decision tree exhibited the best calibration (lowest Brier score = 0.1142). DCA indicated that random forest provided the highest net benefit across most threshold probabilities. Multivariable analysis identified history of ocular disease (OR = 3.529, 95% CI: 2.430-5.139) as the strongest positive predictor, while higher HbA1c, lower self-efficacy, higher perceived severity, absence of exercise therapy, and smoking were associated with poorer adherence.
Conclusion:
Machine learning models demonstrated comparable discriminative performance to traditional logistic regression for predicting DR screening adherence, while offering distinct profiles in calibration and clinical net benefit. Integration of PMT-based psychological constructs with clinical and behavioral factors provides a multidimensional framework for understanding screening behavior and supports risk stratification for precision diabetes care.