Related Experiment Videos
Underperformance of Machine Learning Algorithms Predicting Extended Lengths of Stay and Readmission in
Marium M Raza1, Michelle Riyo Shimizu1, Pengwei Xiao1
1Department of Orthopaedic Surgery, Bioengineering Laboratory, Massachusetts General Hospital, Harvard Medical School, Boston, Massachusetts.
Background:
The demand for total hip arthroplasty (THA) is increasing, yet disparities in access and outcomes persist across racial, ethnic, and socioeconomic groups. Machine learning (ML) models can aid in predicting THA complications such as prolonged lengths of stay (LOS) and 30-day readmission, which is particularly useful for populations at risk of poorer outcomes. However, limited studies to date have assessed ML prediction performance in smaller patient subcohorts that are less commonly represented. Therefore, this study aimed to assess the fairness and performance of ML model prediction of prolonged LOS and 30-day readmission among subcohorts of underrepresented patient groups following primary THA.
Methods:
Using a national database (n = 180,762), ML models were developed to predict prolonged LOS and 30-day readmission post-THA. The model fairness was assessed across demographic (age, sex, race, and ethnicity) and clinical factors (diabetes status). The fairness metrics included equal opportunity, predictive equality, predictive parity, statistical parity, and accuracy equality ratios. Postprocessing and reduction alg orithms were then applied to address underperformance and determine if fairness metrics improved.
Results:
The fairness analysis of both LOS and readmission algorithms revealed lower model performance for Hispanic/Latinx patients, women, and patients who had diabetes across key metrics, including predictive parity and statistical parity. While mitigation algorithms improved ML performance across several fairness metrics, they also resulted in the worsening of other metrics.
Conclusions:
While ML models for THA outcome prediction can show robust overall predictive accuracy, these findings highlight the importance of evaluating model fairness across smaller patient subcohorts. Mitigation algorithms can be useful, but they should be embedded within a broader equity-focused framework prior to clinical integration of ML algorithms.