Related Experiment Video
Updated: Sep 19, 2026

A Periprosthetic Joint Candida albicans Infection Model in Mouse
Published on: February 2, 2024
Prediction Models for Periprosthetic Joint Infection: A Systematic Review of Traditional and Machine Learning
Farzad Pourghazi1, Seyed Mohammad Amin Alavi2, Fabio Borgonovo1,3
1Division of Public Health, Infectious Diseases and Occupational Medicine, Department of Medicine, Mayo Clinic College of Medicine and Science, Mayo Clinic, Rochester, MN.
Objective:
To systematically review the development, validation, and performance of traditional and machine learning (ML) prediction models for periprosthetic joint infection (PJI) risk, diagnosis, and clinical outcomes.
Methods:
We conducted a systematic review according to PRISMA guidelines to identify studies that developed or validated multivariable prediction models for PJI using traditional statistical or ML approaches. Searches were performed on December 10, 2025, across Ovid Embase, MEDLINE, Scopus, Web of Science, ClinicalTrials.gov, and CENTRAL.
Results:
Among 4565 identified records, 34 studies met inclusion criteria. Reported discrimination varied across model types and clinical end points. Among traditional risk-prediction models, 54.5% showed acceptable or good discrimination (0.7 ≤ area under the receiver-operating characteristic curve [AUROC] < 0.8), and 27.3% showed excellent discrimination (0.8 ≤ AUROC < 0.9). Traditional diagnostic and outcome prediction models showed good to outstanding discrimination. Among ML models, 50.0% of risk models showed acceptable or good discrimination (0.7 ≤ AUROC < 0.8), while 60.0% of diagnostic models reported AUROC of 0.9 or more. However, calibration reporting was inconsistent, particularly among ML studies, and external validation was limited. Model performance generally decreased in external validation cohorts compared with derivation cohorts. Substantial heterogeneity among included studies and model designs precluded direct crossmodel comparisons; therefore, differences in AUROC should not be interpreted as evidence of superiority between approaches.
Conclusion:
Prediction models for PJI demonstrate variable performance across risk prediction, diagnosis, and outcome prediction settings. However, substantial clinical and methodological heterogeneity, limited calibration reporting, and inadequate external validation restrict assessment of model reliability and generalizability. Current evidence does not support conclusions regarding the superiority between approaches.

