Related Experiment Video
Updated: May 22, 2025

Author Spotlight: Impact of Intergenic Interactions on Disease-Identifying Dark Biomarkers
Published on: March 1, 2024
Analysis of the 50-mile ultramarathon distance using a predictive XGBoost model
Jonas Turnwald1, David Valero2, Pedro Forte3
1Centre for Rehabilitation and Sports Medicine, University Hospital Bern, Inselspital Bern, University of Bern, Bern, Switzerland.
Abstract:
Although the 50-mile ultramarathon is one of the most common race distances, it has received little scientific attention. The objective of this study was to assess how an athlete's age group, sex, nationality, and the race location, affect race speed. Utilizing a dataset with ultramarathon races from 1863 to 2022, a machine learning model based on the XGBoost algorithm was developed to predict the race speed based on the aforementioned variables. Model explainability tools, including model features relative importances and prediction distribution plots were then used to investigate how each feature affects the predicted race speed. The most important features, with respect to the predictive power of the XGBoost model, were the location of the race and the athlete's gender. The top 3 countries with the fastest predicted median race speeds were Slovenia, New Zealand, and Bulgaria for nationality and New Zealand, Croatia, and Serbia for the race location. The fastest median race speed was predicted for the age group 20-24 years, but a marked age-related performance decline only became apparent from the age group 40-44 years onward. Model predictions for male athletes were faster than for female athletes. This study offers insights into factors influencing race speed in 50-mile ultramarathons, which may be beneficial for athletes, coaches, and race organizers. The identification of nationalities and event countries with fast race speeds provides a foundation for further exploration in the field of ultramarathon events.
More Related Videos
Related Concept Videos
End Point Prediction: Gran Plot
For potentiometric titration, the Gran plot is created by plotting...
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Quantifying and Rejecting Outliers: The Grubbs Test
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Regression Toward the Mean

