Related Experiment Videos
A deep-learning approach to predict mortality in very-low-birth-weight infants treated in the intensive care units
Nishankul Bozhbanbayeva1, Arailym Abilbayeva2, Anel Tarabayeva2
1Neonatology Department, Asfendiyarov Kazakh National Medical University, Almaty, Kazakhstan.
Background And Objectives:
Predicting mortality in very low-birth-weight (VLBW) infants remains a critical challenge in neonatology. While traditional scoring systems and standard machine learning models are widely used, they often fail to capture the complex, non-linear interactions and time-dependent dynamics of neonatal clinical courses. This study addresses these limitations by evaluating DeepSurv to enhance predictive accuracy and clinical risk stratification in the neonatal intensive care unit. This study aimed to thoroughly compare the performance of the DeepSurv model with that of traditional machine learning models, such as Cox proportional hazards (CoxPH) and random survival forest (RSF), including temporal metrics.
Materials And Methods:
A retrospective analysis of 958 VLBW neonates was performed using data from an integrated health information system from January 2021 to May 2025. The Boruta algorithm, a feature selection method that identifies statistically significant predictors by comparing their importance scores with those of random, shadow variables, identified 11 such predictors in our study.
Results:
DeepSurv demonstrated superior discriminatory ability with an overall AUC of 0.95, substantially outperforming RSF (AUC of 0.89) and CoxPH (AUC of 0.86). The temporal AUC analysis also favored DeepSurv. Moreover, DeepSurv recorded the lowest Brier scores (IBS 0.071) throughout the entire follow-up period compared with CoxPH (IBS 0.129) and RSF (IBS 0.126). Calibration curves showed acceptable agreement between predicted and observed risk by day 28 for all models, whereas at day 7 calibration was weaker, particularly for DeepSurv at low predicted-risk values, and Kaplan-Meier analysis, along with log-rank tests, demonstrated their effectiveness in risk stratification. Decision curve analysis revealed that the DeepSurv model had the highest clinical utility at both day 7 and day 28, underscoring its potential impact on neonatal care.
Conclusion:
DeepSurv outperformed traditional models in predicting VLBW mortality, demonstrating superior clinical utility and robustness against overfitting. While the early neonatal period (days 5-7) remains the most challenging to predict, DeepSurv provides a powerful framework for personalized risk stratification, though further optimization is needed to capture the highly dynamic changes of the first week of life.