Performance Drift in Machine Learning Models for Cardiac Surgery Risk Prediction: Retrospective Analysis.
Tim Dong1, Shubhra Sinha1, Ben Zhai2
1Bristol Heart Institute, Translational Health Sciences, University of Bristol, Bristol, United Kingdom.
Jmirx Med
|June 18, 2024
Summary
Machine learning models for cardiac surgery risk prediction show performance decline over time due to dataset drift. Extreme gradient boosting and random forest models outperformed traditional scores like EuroSCORE II.
Area of Science:
- Cardiovascular Surgery
- Medical Informatics
- Machine Learning in Healthcare
Background:
- Traditional risk scores (e.g., EuroSCORE II) for adult cardiac surgery mortality are prone to miscalibration and poor generalization.
- Dataset drift, where deployed ML models encounter data different from their training data, hinders ML adoption in clinical practice.
- Understanding performance drift in ML models is crucial for their reliable application in real-world healthcare settings.
Purpose of the Study:
- To assess the extent of performance drift in machine learning (ML) models for cardiac surgery risk prediction over time.
- To investigate the influence of dataset drift and variable importance drift on model performance.
- To compare the performance drift of ML models against the EuroSCORE II risk score.
Main Methods:
- Retrospective analysis of prospectively collected UK cardiac surgery data (2012-2019) from 227,087 adult patients.
- Development and temporal validation of five novel ML mortality prediction models alongside EuroSCORE II.
- Assessment of relationships between variable importance drift, performance drift, and dataset drift using a consensus metric.
Main Results:
- A significant decrease in overall performance was observed across all models (P<.0001).
- Extreme gradient boosting (CEM 0.728) and random forest (CEM 0.727) demonstrated the best performance.
- EuroSCORE II exhibited the poorest performance; sharp changes in variable importance and dataset drift correlated with performance decreases.
Conclusions:
- All evaluated models experienced performance degradation over time, highlighting the impact of dataset drift.
- Variable importance drift detection underscores limitations of logistic regression and the effects of dataset drift in cardiac surgery risk prediction.
- Further research is needed to explore ensemble ML models for potentially improved performance and robustness.


