Performance, Heterogeneity, and Methodological Quality of YOLO-Based Models for Fracture Detection: A Systematic
Zexi Wang1, Yuan Zhang2, Yixi Wang1
1Department of Minimally Invasive Spine and Precision Orthopedics, The First Affiliated Hospital of Xinjiang Medical University, Uygur Autonomous Region, Urumqi, 830054, Xinjiang, China.
Abstract:
This systematic review and meta-analysis aimed to evaluate the technical detection performance of YOLO-based models for fracture detection and examine clinical and methodological factors associated with performance heterogeneity across studies. We conducted a PRISMA-compliant systematic review and meta-analysis registered in PROSPERO. PubMed and Web of Science were searched from January 1, 2018, to December 15, 2025, with backward citation tracking. Study quality was assessed using QUADAS-AI. Random-effects meta-analysis with restricted maximum likelihood estimation and Hartung-Knapp-Sidik-Jonkman adjustment was used to pool mAP@0.5. Twenty-nine studies were included. The pooled mAP@0.5 was 0.797 (95% CI 0.725-0.854), with substantial heterogeneity (I2 = 99.2%) and a 95% prediction interval of 0.333-0.969. Population was the only prespecified subgroup factor that remained statistically significant, with lower pooled performance in pediatric than adult cohorts; however, only 7 pediatric studies were available, so this finding was considered preliminary. Anatomical-site differences were non-significant after adjustment, and YOLO version was not significantly associated with performance. A risk-of-bias sensitivity analysis excluding 12 high-risk studies yielded a pooled mAP@0.5 of 0.830 (95% CI 0.745-0.890). An equal-study-weight sensitivity analysis gave a similar estimate (0.823, 95% CI 0.764-0.870). YOLO-based models showed promising overall technical performance for fracture detection, but the literature was highly heterogeneous. The wide prediction interval indicates that performance may vary substantially across settings, and the pediatric subgroup finding remains preliminary. No statistically significant association between YOLO version and performance was detected. At present, these systems seem better suited to focused decision-support roles than to fully autonomous diagnosis.

