Related Experiment Videos
External Validation of Clinical Risk Scores and Machine Learning Models for Predicting 30-Day Cardiovascular Risk
Aslan Erdoğan1, Şeyma Yeşil1, Gamze Gençol Akçay1
1Department of Cardiology, Cam and Sakura Training and Research Hospital, 34480 Istanbul, Turkey.
Abstract:
Background: Machine learning (ML) has emerged as a promising approach for preoperative cardiovascular risk prediction; however, the generalizability of ML models across institutions remains uncertain. Moreover, comprehensive head-to-head comparisons between ML algorithms and established clinical risk scores for predicting 30-day major adverse cardiac events (MACE) are scarce. We therefore evaluated the performance and external transportability of multiple ML models across independent centers and compared their predictive accuracy with validated benchmark clinical risk scores. Methods: In a site-separated two-center cohort (derivation n = 707, 27 MACE; external validation n = 378, 38 MACE), ten algorithms trained on preoperative variables were externally validated without refitting and benchmarked against the American university of Beirut-HAS2 (AUB-HAS2), American society of anesthesiologists (ASA), and revised cardiac risk index (RCRI). We assessed AUROC, calibration, Brier score, and decision-curve net benefit, with paired bootstrap comparisons, DeLong testing, IDI, and NRI. Results: Among the ML models, no single algorithm consistently outperformed the others across all performance metrics. Naive Bayes achieved the highest external discrimination (AUROC 0.738, 95% CI 0.668-0.804) but showed poor calibration and threshold-dependent clinical utility. Gradient Boosting showed the most favorable balance of discrimination and calibration slope (AUROC 0.707; calibration slope 0.991), although absolute risk remained underestimated in external validation, whereas HistGradient Boosting yielded the best overall probability prediction, with the lowest Brier score (0.087) and the greatest decision-curve net benefit at clinically relevant risk thresholds. However, in external validation, no ML model demonstrated statistically significant AUROC superiority over the AUB-HAS2 score (all p > 0.05), although several models modestly outperformed the RCRI. Conclusions: In this site-separated external validation, ML models showed metric-dependent performance but no discrimination advantage over the AUB-HAS2 index. Given low event counts, flexible-model results are hypothesis-generating. These findings provide a cautionary, reproducible benchmark; local recalibration and prospective evaluation are prerequisites before clinical deployment.