Related Experiment Videos

Underperformance of Machine Learning Models Predicting Readmission and Prolonged Length of Stay Following Total Knee

Michelle R Shimizu1, Marium Raza1, Pengwei Xiao1

  • 1Department of Orthopaedic Surgery, Bioengineering Laboratory, Massachusetts General Hospital, Harvard Medical School, Boston, Massachusetts.

Abstract

Insights

Machine learning models show bias in predicting total knee arthroplasty outcomes for minority groups. Mitigation strategies improved fairness but did not eliminate bias, highlighting the need for equitable AI in healthcare.

Area of Science:

  • Health Informatics
  • Artificial Intelligence in Medicine
  • Health Equity Research

Background:

  • Machine learning (ML) models exhibit high predictive accuracy but underperform for minority subcohorts, potentially worsening health disparities.
  • A "one-size-fits-all" approach in ML decision-making perpetuates inequities for underrepresented patient groups.
  • This study investigates ML model fairness for total knee arthroplasty (TKA) outcomes and explores bias mitigation.

Purpose of the Study:

  • To assess the fairness of validated ML models in predicting TKA outcomes.
  • To explore bias mitigation strategies to enhance equitable prediction performance in diverse patient populations.
  • To identify and address performance disparities in ML models across different demographic groups.

Main Methods:

  • Developed and validated four ML models to predict TKA readmission and prolonged length of stay (pLOS).
  • Performed bias assessment using protected attributes (age, sex, race, ethnicity) and evaluated fairness metrics.
  • Incorporated and trialed three bias mitigation strategies, integrating the most effective for each attribute and reassessing fairness.

Main Results:

  • The random forest model showed the best predictive performance (Readmission AUC = 0.98; pLOS AUC = 0.92).
  • Inferior predictive equality was observed for readmission in women (0.47 vs. men), non-White (0.75 vs. White), and Hispanic/LatinoX (0.53 vs. not Hispanic/LatinoX) subcohorts.
  • Significant disparities in pLOS prediction accuracy were found across underprivileged cohorts, with mitigation strategies showing trade-offs between fairness metrics.

Conclusions:

  • Validated ML models demonstrate "unfairness" or underperformance in predicting TKA readmission and pLOS for smaller subcohorts.
  • Mitigation techniques improved fairness metrics but did not fully eliminate bias, necessitating careful consideration of strategy trade-offs.
  • Substantial efforts are required to correct subcohort bias in ML algorithms before clinical use to ensure patient equity.

Related Concept Videos