Multidisciplinary prediction of running-related injuries using machine learning
Han Wu1, Katherine Brooke-Wavell2, Michael R Barnes3
1School of Sport, Exercise and Health Sciences, Loughborough University, Loughborough, UK. h.wu4@lboro.ac.uk.
NPJ Digital Medicine
|February 6, 2026
Summary
This study developed a machine learning dataset for predicting endurance running injuries using diverse risk factors. The models showed moderate improvement in injury prediction accuracy, with Random Forest performing best.
Area of Science:
- Sports Medicine
- Biomechanical Engineering
- Data Science
Background:
- Endurance running-related injuries (RRIs) have complex, multifactorial causes.
- Existing research often overlooks multidisciplinary risk factors for individualized RRI prediction.
Purpose of the Study:
- To create a machine learning-ready dataset for weekly RRI prediction.
- To evaluate the efficacy of machine learning models using multidisciplinary risk factors.
Main Methods:
- Collected data on genetic, historical, biomechanical, physiological, and training factors from 142 competitive runners over 12 months.
- Developed and tested machine learning models using both high-evidence and broader sets of risk factors.
- Prospectively monitored runners for RRIs, accumulating 6181 weekly samples.
Main Results:
- Machine learning models achieved an AUC of 0.784 ± 0.014, showing moderate improvement over previous RRI prediction.
- Random Forest models demonstrated the highest performance (AUC = 0.781 ± 0.016).
- Logistic regression performance significantly improved with a broader set of risk factors.
Conclusions:
- Introduced a reproducible framework for machine learning-based sports injury prediction.
- Provided a valuable dataset for future large-scale sports injury analytics.
- Highlighted the importance of multidisciplinary data and model selection for accurate RRI prediction.
More Related Videos
Related Concept Videos
Wald-Wolfowitz Runs Test II
558
The Wald-Wolfowitz runs test, commonly referred to as the runs test, is a nonparametric test used to assess the randomness of ordered data. The test evaluates the number of runs, which are consecutive sequences of similar elements within the data. If the number of runs is significantly higher or lower than expected, the data is considered non-random, indicating a detectable pattern or structure.
For binary data, runs are identified using symbols such as + and −, or equivalently, 1s and 0s. In...
For binary data, runs are identified using symbols such as + and −, or equivalently, 1s and 0s. In...
558
Wald-Wolfowitz Runs Test I
964
The Wald-Wolfowitz test, also known as the runs test, is a nonparametric statistical test used to assess the randomness of a sequence of two different types of elements (e.g., positive/negative values, successes/failures). It examines whether the order of the elements in a sequence is random or if there is a pattern or trend present. This nonparametric test applies to any ordered data despite the population and sample data distribution, even if a higher sample size is available.
The test works...
The test works...
964
Predicting Molecular Geometry
46.0K
VSEPR Theory for Determination of Electron Pair Geometries
46.0K
Machines
581
Machines are complex structures consisting of movable, pin-connected multi-force members that work together to transmit forces. One example of a machine is the cutting plier, which is used to cut wires by applying forces to its handles. When equal and opposite forces are exerted on the handles of the cutting plier, they cause the cutting edges to come together and apply equal and opposite reaction forces on the wire, which are greater than the applied forces.
A free-body diagram of the...
A free-body diagram of the...
581
Machines: Problem Solving II
673
Machines are complex structures consisting of movable, pin-connected multi-force members that work together to transmit forces. Consider a lifting tong carrying a 100 kg load. It comprises movable sections DAF and CBG linked together with member AB.
673
Prediction Intervals
3.4K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
3.4K


