Related Experiment Video
Updated: Aug 28, 2026

An Inertial Measurement Unit Based Method to Estimate Hip and Knee Joint Kinematics in Team Sport Athletes on the Field
Published on: May 26, 2020
Machine Learning Model Development and Evaluation for Non-contact Lower Limb Injury Risk Prediction in Elite Male
Chris Leckey1,2, Nicol van Dyk3,4, Jakim Berndsen3
1School of Public Health, Physiotherapy & Sports Science, University College Dublin, Dublin, Ireland. christopher.leckey@ucdconnect.ie.
Background:
The prevalence, incidence rate and burden of injuries are high in rugby union. Despite the growing application of machine learning for injury risk analysis in sport, its use in rugby union has been primarily confined to single-team and single-season study cohorts. The aim of this study was to evaluate the performance of machine learning-based injury risk prediction models in elite male rugby union and to interpret their potential applicability in professional sporting environments.
Methods:
Multi-season data (2021/22-2023/24) relating to 129 professional male rugby union players from the four Irish provincial teams (Connacht Rugby, Leinster Rugby, Munster Rugby and Ulster Rugby), along with the Ireland men's national team, were included in the analyses. External workload data from Global Navigation Satellite System (GNSS) units, self-perceived wellness scores, and musculoskeletal screening measures were modelled using logistic regression, Support Vector Machine (SVM), Random Forest, Extreme Gradient Boosting (XGBoost), and Categorical Boosting (CatBoost). 'Injury' was defined as a non-contact, soft-tissue injury to the lower limb, with incidents logged by team medical staff in a centralised athlete management system.
Results:
CatBoost achieved the highest values of area under the receiver operating characteristic curve (ROC AUC = 0.66), area under the precision-recall curve (AUPRC = 0.008, prevalence = 0.0039), and precision (0.013).
Conclusion:
Predictive performance across the models was poor, with low precision scores highlighting each model's inability to effectively identify impending injuries. The resulting high false-positive rate (1 correct prediction for every 77 false predictions) underscores the limited applicability of the tested machine learning approaches within the modelling framework analysed in this study.