使用预测XGBoost模型对50英里超级马拉松距离的分析
Jonas Turnwald1, David Valero2, Pedro Forte3
1Centre for Rehabilitation and Sports Medicine, University Hospital Bern, Inselspital Bern, University of Bern, Bern, Switzerland.
Scientific reports
|March 16, 2025
概括
比赛地点和运动员的性别显著影响50英里超级马拉松速度. 性能随着年龄的增长而下降,从40-44岁开始,在特定的国籍和比赛地点发现的速度更快.
科学领域:
- 运动科学 运动科学 运动科学
- 耐力运动研究 耐力运动研究
- 人类绩效分析 人类绩效分析
背景情况:
- 50英里超级马拉松很受欢迎,但研究不足.
- 影响超级马拉松比赛速度的因素需要科学研究.
研究的目的:
- 分析运动员人口统计 (年龄组,性别,国籍) 和比赛地点对50英里超级马拉松速度的影响.
- 使用机器学习开发一个用于赛车速度的预测模型.
主要方法:
- 利用了50英里超级马拉松比赛 (1863-2022) 的综合数据集.
- 开发并应用XGBoost机器学习模型用于速度预测.
- 采用模型可解释性工具来识别关键影响因素.
主要成果:
- 比赛地点和运动员的性别是比赛速度的最重要的预测因素.
- 最快的中位速度预测来自斯洛文尼亚,新西兰和保加利亚的运动员.
- 从40-44岁的年龄组开始,表现下降变得明显;男性预测的速度比女性更快.
结论:
- 确定了影响50英里超级马拉松表现的关键人口和地理因素.
- 调查结果为运动员,教练和比赛组织者提供了宝贵的见解.
- 为未来研究超级马拉松赛事动态和性能优化提供了基础.
相关概念视频
End Point Prediction: Gran Plot
218
A Gran plot is used to predict the equivalence volume or endpoint of a potentiometric or acid-base titration without reaching the endpoint. Typically, titration data is collected as a function of the titrant's volume up to a point less than the equivalence volume and then transformed into a linear format. The straight line is extended to the x-axis, indicating the necessary titrant volume to achieve the equivalence point.
For potentiometric titration, the Gran plot is created by plotting...
For potentiometric titration, the Gran plot is created by plotting...
218
Prediction Intervals
2.2K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.2K
Quantifying and Rejecting Outliers: The Grubbs Test
1.4K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
1.4K
Regression Analysis
5.5K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
5.5K
Multiple Regression
2.9K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
2.9K
Regression Toward the Mean
6.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.3K


