Related Experiment Videos
Underperformance of Machine Learning Models Predicting Readmission and Prolonged Length of Stay Following Total Knee
Michelle R Shimizu1, Marium Raza1, Pengwei Xiao1
1Department of Orthopaedic Surgery, Bioengineering Laboratory, Massachusetts General Hospital, Harvard Medical School, Boston, Massachusetts.
Background:
While machine learning (ML) models demonstrate high predictive accuracy, recent studies reveal that ML models underperform for smaller subcohorts such as racial and ethnic minorities, suggesting inherent ML biases that may exacerbate health disparities. To assume a "one-size-fits-all" approach perpetuates inequities in decision-making for underrepresented groups. This study assessed validated ML model "fairness" for total knee arthroplasty (TKA) outcomes and explored bias mitigation strategies to enhance equitable prediction performance.
Methods:
There were four ML models that were developed and validated to predict readmission and prolonged lengths of stay (pLOSs) following TKA. Bias assessment was performed using protected attributes (age, sex, race, and ethnicity), and various fairness metrics were evaluated. There were three bias mitigation strategies incorporated and trialed for each algorithm, and the most effective strategy for a given protected attribute was integrated with the original model and reassessed based on the fairness metrics.
Results:
The random forest model had the best predictive performance (ReadmissionAUC = 0.98; pLOSAUC = 0.92). Inferior predictive equality for readmission was demonstrated in women (0.47 versus men), non-White (0.75 versus White), and Hispanic or LatinoX (0.53 versus not Hispanic/LatinoX) subcohorts. Significant differences in all, but accuracy equality metrics for pLOS were unveiled across underprivileged cohorts. Mitigation strategies were effective in both ML models; however, trade-offs were observed between the fairness metrics.
Conclusions:
Our study highlights the "unfairness" or underperformance of current validated ML models in predicting readmission and pLOS in smaller subcohorts after TKA. While mitigation techniques improved fairness metrics, no approach fully eliminated bias, underscoring the need to carefully consider the trade-offs of each strategy. Our findings suggest substantial efforts should be made to correct potential bias in subcohorts before clinical algorithms are published or utilized as clinical decision-support tools to ensure patient equity.
Insights
Machine learning models show bias in predicting total knee arthroplasty outcomes for minority groups. Mitigation strategies improved fairness but did not eliminate bias, highlighting the need for equitable AI in healthcare.
Area of Science:
- Health Informatics
- Artificial Intelligence in Medicine
- Health Equity Research
Background:
- Machine learning (ML) models exhibit high predictive accuracy but underperform for minority subcohorts, potentially worsening health disparities.
- A "one-size-fits-all" approach in ML decision-making perpetuates inequities for underrepresented patient groups.
- This study investigates ML model fairness for total knee arthroplasty (TKA) outcomes and explores bias mitigation.
Purpose of the Study:
- To assess the fairness of validated ML models in predicting TKA outcomes.
- To explore bias mitigation strategies to enhance equitable prediction performance in diverse patient populations.
- To identify and address performance disparities in ML models across different demographic groups.
Main Methods:
- Developed and validated four ML models to predict TKA readmission and prolonged length of stay (pLOS).
- Performed bias assessment using protected attributes (age, sex, race, ethnicity) and evaluated fairness metrics.
- Incorporated and trialed three bias mitigation strategies, integrating the most effective for each attribute and reassessing fairness.
Main Results:
- The random forest model showed the best predictive performance (Readmission AUC = 0.98; pLOS AUC = 0.92).
- Inferior predictive equality was observed for readmission in women (0.47 vs. men), non-White (0.75 vs. White), and Hispanic/LatinoX (0.53 vs. not Hispanic/LatinoX) subcohorts.
- Significant disparities in pLOS prediction accuracy were found across underprivileged cohorts, with mitigation strategies showing trade-offs between fairness metrics.
Conclusions:
- Validated ML models demonstrate "unfairness" or underperformance in predicting TKA readmission and pLOS for smaller subcohorts.
- Mitigation techniques improved fairness metrics but did not fully eliminate bias, necessitating careful consideration of strategy trade-offs.
- Substantial efforts are required to correct subcohort bias in ML algorithms before clinical use to ensure patient equity.