Related Experiment Video
Updated: Aug 10, 2026

Biomechanical Changes Related to Low Back Pain: An Innovative Tool for Movement Pattern Assessment and Treatment Evaluation in Rehabilitation
Published on: December 13, 2024
Incomplete reporting persists in orthopedic machine learning models: a systematic review of TRIPOD and TRIPOD+AI
R Harmen Kuijten1, Tom M de Groot2, Maarten A van Weezenbeek3
1Division of Imaging and Oncology, University Medical Center Utrecht, Heidelberglaan 100, Utrecht 3584 CX, The Netherlands.
Background And Objectives:
Machine learning (ML) prognostic models in orthopedic surgery are published at an accelerating pace, yet whether reporting meets the Transparent Reporting of a multivariable prediction model for Individual Prognosis Or Diagnosis (TRIPOD) + artificial intelligence (AI; 2024) transparency standards remains unclear. We examined (1) characteristics and trends of preoperative ML prognostic models for perioperative outcomes through December 31, 2024; (2) TRIPOD+AI reporting completeness among the 25 highest-impact orthopedic journals; and (3) how derived TRIPOD 2015 completeness compares with prior systematic reviews.
Methods:
We performed a systematic review of studies published through December 31, 2024, developing or externally evaluating preoperative ML prognostic models predicting any intra- or post-operative outcome in orthopedic surgery. Study characteristics and trends were extracted from 433 eligible studies. TRIPOD+AI reporting completeness was assessed in a subset of 93 studies (80 development-applicable and 20 evaluation-applicable; not mutually exclusive as 7 studies performed both) published in the 25 highest-impact orthopedic journals. TRIPOD 2015 reporting completeness was derived from TRIPOD+AI items.
Results:
Nearly half of all studies were published in 2023-2024, a total of 27% of development studies had high sample size concern, 36% reported calibration, and only 13% provided an accessible model. The median TRIPOD+AI completeness was 45% (IQR 40-50). The median TRIPOD 2015 completeness was similar in development studies (54% [IQR 28-84] vs 50% [IQR 27-79]) but improved in evaluation studies (80% [IQR 45-95] vs 61% [IQR 43-90]) compared with prior systematic reviews.
Conclusion:
Despite rapid growth with more than 400 ML orthopedic prognostic prediction models, this baseline assessment reveals substantial gaps in reporting completeness. There is considerable room for improvement, particularly in sample size justification and performance reporting. Future work should prioritize TRIPOD+AI implementation, journal enforcement, and periodic reevaluation of reporting completeness as post-2024 publications mature.