Related Experiment Video
Updated: Aug 14, 2026

08:26
Comparative Analysis of Lower Limb Kinematics between the Initial and Terminal Phase of 5km Treadmill Running
Published on: July 17, 2020
Using Machine Learning to Explain Recreational Running Injury: An Exploratory Classification Analysis
Bradley S Neal1,2, Stella Hadjiantoni3, Allison H Gruber4
1Sports and Exercise Medicine, Queen Mary University of London, London, United Kingdom.
Journal of Sport Rehabilitation
|August 12, 2026
Summary
Machine learning models can predict running injuries by analyzing run data. Critical power emerged as the most significant factor, indicating higher values increase injury risk.
Area of Science:
- Sports Medicine
- Biomechanical Engineering
- Data Science
Background:
- Running is beneficial for health but carries an injury risk that is not fully understood.
- Wearable technology and machine learning (ML) offer promising tools to address this challenge.
- This study aimed to identify key factors, optimal ML models, and sample sizes for future research on running injuries.
Purpose of the Study:
- To identify critical factors predicting running-related injuries using ML.
- To determine the best-performing ML model for injury risk classification.
- To estimate the required sample size for future studies on running injury development.
Main Methods:
- An exploratory classification analysis was conducted using data from 88 recreational runners over 12 weeks.
- Heart rate, inertial measurement unit, and GPS data were collected from wristwatches.
- Seven ML models were evaluated using 10-fold cross-validation, with performance metrics including accuracy, precision, and recall. SHapley Additive exPlanations identified factor importance.
Main Results:
- The dataset comprised 4758 run instances from 80 participants.
- Four ML models (CatBoost, GMB, Light Gradient Boosting Machine, GXBoost) demonstrated high performance (>0.90 average).
- CatBoost achieved 96% accuracy in classifying injured runs, with critical power identified as the most significant predictor.
Conclusions:
- Boosting ML models, particularly CatBoost, are recommended for future studies on running injury development.
- Critical power should be included as a key factor in predictive models.
- Results should not be interpreted as causal, and model performance may be overestimated.