Related Experiment Videos
Advancing ST-elevated myocardial infarction mortality risk prediction in Asian populations through explainable and
Sazzli Kasim1,2, Lim Bing Feng3, Putri Nur Fatin Amir Rudin3
1Cardiovascular Advancement and Research Excellence Institute (CARE Institute), Universiti Teknologi MARA, Selangor, Malaysia.
Insights
Explainable machine learning models significantly outperform traditional risk scores for predicting in-hospital mortality in Asian ST-segment elevation myocardial infarction (STEMI) patients, improving risk stratification and patient outcomes.
Area of Science:
- Cardiovascular Medicine
- Artificial Intelligence in Healthcare
- Biostatistics
Background:
- Traditional risk scores for ST-segment elevation myocardial infarction (STEMI), like the Thrombolysis in Myocardial Infarction (TIMI) score, show suboptimal performance in Asian populations due to genetic and clinical differences.
- Accurate prediction of in-hospital mortality is crucial for managing high-risk STEMI patients.
- Existing predictive models often lack crucial explainability and probability calibration, limiting their clinical utility.
Purpose of the Study:
- To develop and validate explainable, well-calibrated machine learning (ML) models for predicting in-hospital mortality in Asian STEMI patients.
- To benchmark the performance of these ML models against the established TIMI risk score.
- To enhance model interpretability using SHAP (SHapley Additive exPlanations) analysis and systematically address probability calibration.
Main Methods:
- A retrospective cohort study of 49,574 Asian STEMI patients from the Malaysian National Cardiovascular Disease registry (2006-2021).
- Temporal data splitting for training (2006-2018), calibration (2019), and independent testing (2020-2021).
- Development and comparison of ML algorithms (logistic regression, SVM, random forest, GBM, XGBoost), including stacked ensembles, with performance evaluated using AUC-ROC, accuracy, recall, specificity, Brier score, and NRI. Isotonic regression for calibration and SHAP for interpretability.
Main Results:
- The calibrated logistic regression (LR) model demonstrated superior performance (AUC: 0.8884, Brier score: 0.0598, NRI vs. TIMI: 0.5828).
- SHAP analysis confirmed model predictions aligned with clinical reasoning, enhancing interpretability.
- Probability calibration significantly improved model reliability, indicated by a lower Brier score.
Conclusions:
- Calibrated logistic regression models, enhanced with SHAP explainability and robust calibration, significantly outperform the TIMI score in predicting in-hospital mortality for Asian STEMI patients.
- This approach enhances predictive accuracy, reliability, and interpretability, facilitating personalized risk stratification.
- The findings support the potential for clinical integration of these advanced models to improve patient outcomes in diverse Asian populations.
Background:
Traditional risk scores for ST-segment elevation myocardial infarction (STEMI), such as the Thrombolysis in Myocardial Infarction (TIMI) score, were developed predominantly in Western populations and may exhibit suboptimal performance in Asian patients due to differing clinical and genetic profiles. Accurate in-hospital mortality prediction is essential for optimizing clinical management in this high-risk group. However, key aspects such as model explainability (e.g., SHapley Additive exPlanations (SHAP) analysis) and probability calibration are often neglected, limiting the clinical utility and trustworthiness of predictive models.
Objectives:
This study aimed to develop and validate explainable, well-calibrated machine learning models for predicting in-hospital mortality among Asian STEMI patients, benchmarking their performance against the TIMI risk score. Interpretability was enhanced using SHAP for both global and local explanations, and model calibration was systematically addressed.
Methods:
We conducted a retrospective cohort study using data from 49,574 Asian STEMI patients in the Malaysian National Cardiovascular Disease registry (2006-2021). A temporal split was applied to simulate prospective deployment: data from 2006-2018 were used for model training, 2019 data for calibration, and 2020-2021 data as an independent test set. Multiple ML algorithms including logistic regression (LR), support vector machine, random forest, gradient boosting machine, and XGBoost were developed and compared. Stacked ensemble models were constructed using combinations of these base learners. Performance was evaluated on the independent test set using area under the receiver operating characteristic curve (AUC-ROC), accuracy, recall, specificity, Brier score (for calibration), and Net Reclassification Index (NRI), with benchmarking against the TIMI score. Isotonic regression was applied for probability calibration, and SHAP was used for interpretability.
Results:
The calibrated LR model achieved the best overall performance (AUC = 0.8884, 95% CI (confidence interval): 0.8756-0.9011; accuracy = 0.8538; recall = 0.7746; specificity = 0.8617; Brier score = 0.0598; NRI = 0.5828 vs. TIMI). SHAP analysis confirmed that model predictions were aligned with established clinical reasoning, enhancing interpretability. Probability calibration further improved model reliability, as evidenced by a reduced Brier score.
Conclusions:
Calibrated LR, supported by SHAP-based explainability and robust probability calibration, significantly outperformed the TIMI score for in-hospital mortality prediction in Asian STEMI patients. This approach improves predictive accuracy, reliability, and interpretability, supporting more personalized and clinically trustworthy risk stratification. These findings highlight strong potential for real-world clinical integration and improved patient outcomes in diverse Asian populations.