Related Experiment Videos
Development and Validation of a Machine Learning Model for Predicting 28-Day Mortality in Patients with
Hai-Peng Wu1, Kai Liu2, Xiao-Yi Hu1
1Department of Anesthesiology, the Second Affiliated Hospital of Nanjing Medical University, Nanjing, China.
Abstract:
IntroductionPatients with sepsis-associated encephalopathy (SAE) and heart failure (HF) represent a clinically relevant but understudied subgroup in the intensive care unit, and early risk stratification may be important for guiding management. This study aimed to develop and temporally validate machine-learning models for predicting 28-day mortality in patients with SAE and HF.MethodsIn this retrospective cohort study, patients with SAE and HF were identified from the Medical Information Mart for Intensive Care IV (MIMIC-IV) and III (MIMIC-III) databases. The MIMIC-IV cohort was randomly divided into a training set and an internal testing set, whereas the MIMIC-III cohort served as a temporally distinct testing set. Variables with more than 30% missing data were excluded, and the remaining missing values were handled using multiple imputation. Candidate predictors were first screened by univariable analysis and then entered into multivariable logistic regression to identify independent predictors of 28-day mortality. Six variables (age, weight, respiratory rate, Logistic Organ Dysfunction System score, first 24-h urine output, and Charlson Comorbidity Index) were used to construct five models: logistic regression, XGBoost, LightGBM, support vector machine, and AdaBoost. Model performance was evaluated using the area under the receiver operating characteristic curve (AUROC), classification metrics, calibration, and decision-curve analysis.ResultsA total of 1249 patients with SAE and HF were included, comprising 833 in the training set, 205 in the internal testing set, and 211 in the temporal testing set. Machine-learning models and logistic regression outperformed conventional severity scores (SOFA, SAPS II, and LODS). In the temporal testing set, XGBoost achieved the highest AUROC (0.824), whereas logistic regression showed similar discrimination (AUROC 0.816) but more stable calibration and better overall balance across performance domains.ConclusionWe developed and temporally validated models for predicting 28-day mortality in patients with SAE and HF. Among the evaluated models, logistic regression showed the most balanced overall performance and may be the most clinically applicable model for individualized risk stratification.