在中国职业足球联赛中,顶级和二级联赛之间的胜利决定因素:可解释的机器学习方法
Bo Yuan1, Jiaxuan Zhu1, Pengyu Pan2
1College of Physical Education and Sport, Beijing Normal University, Beijing, China.
BMC sports science, medicine & rehabilitation
|April 21, 2025
概括
在中国职业足球中获胜取决于精确的射击和有效的防守. 在盒子内射击目标 (SOTIB) 和空隙是关键,在CFACL中犯规更为关键. 教练可以根据这些见解来定制战术.
科学领域:
- 运动科学 运动科学 运动科学
- 足球分析 足球分析
- 绩效分析 绩效分析
背景情况:
- 了解比赛结果的决定因素对于职业足球的战术发展至关重要.
- 之前的研究已经探讨了各种绩效指标,但联盟特定的分析是必不可少的.
研究的目的:
- 研究和比较影响中国两个职业足球联赛比赛结果的关键因素:中国超级足球联赛 (CSL) 和中国足球协会中国联赛 (CFACL).
- 通过使用先进的机器学习技术,确定最具影响力的绩效指标,用于预测每个联盟的胜利.
主要方法:
- 利用极端梯度提升 (XGBoost) 来分析2017-2019赛季的1440场比赛.
- 使用夏普利添加剂解释 (SHAP) 来解释25个重要指标的变量重要性.
- 专注于得分和防守指标,包括射击目标在盒子内 (SOTIB),射击,射击目标 (SOT),清空和犯规.
主要成果:
- 评分绩效指标,特别是SOTIB,是CSL和CFACL比赛结果的最重要的决定因素.
- 防守排名第二的重要,突出了他们在比赛成功中的关键作用.
- 与CSL相比,防守犯规对CFACL的结果产生了更大的影响.
- 在罚球区内射击的精确性对这两个联盟来说比射击量更为关键.
结论:
- 教练应该优先考虑射门准确性,而不是射门频率,以提高获胜的概率.
- 对于CFACL球队来说,当高质量的传球具有挑战性时,设置的棋子提供了一个可行的替代攻击策略.
- 这些发现为在中国职业足球中开发联盟特定的战术策略提供了可操作的见解.
相关概念视频
Multiple Regression
2.9K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
2.9K
Classification of Systems-II
119
Continuous-time systems have continuous input and output signals, with time measured continuously. These systems are generally defined by differential or algebraic equations. For instance, in an RC circuit, the relationship between input and output voltage is expressed through a differential equation derived from Ohm's law and the capacitor relation,
119
Standard Deviation
15.6K
The most commonly used measure of variation is the standard deviation. It is a numerical value measuring how far data values are from their mean. The standard deviation value is small when the data are concentrated close to the mean, exhibiting slight variation or spread. The standard deviation value is never negative, it is either positive or zero. The standard deviation is larger when the data values are more spread out from the mean, which means the data values are exhibiting more variation.
15.6K
Aggregates Classification
289
Aggregate classification is generally based on its size, petrographic characteristics, weight, and source. Size classification ranges from coarse to fine aggregates, defined by the size of the particles. Coarse aggregates are particles that do not pass through ASTM sieve No. 4, and aggregates that pass through the sieve are fine aggregates.
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
289
Regression Analysis
5.5K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
5.5K
Outliers and Influential Points
3.9K
An outlier is an observation of data that does not fit the rest of the data. It is sometimes called an extreme value. When you graph an outlier, it will appear not to fit the pattern of the graph. Some outliers are due to mistakes (for example, writing down 50 instead of 500), while others may indicate that something unusual is happening. Outliers are present far from the least squares line in the vertical direction. They have large "errors," where the "error" or residual is the...
3.9K


