可解释的机器学习用于评估高速公路撞车严重程度的风险因素
Seyed Alireza Samerei1, Kayvan Aghabayk1
1School of Civil Engineering, College of Engineering, University of Tehran, Tehran, Iran.
概括
可解释机器学习揭示了影响碰撞严重性的因素. 轻量化交通,大型卡车,高速,年轻司机和夜间驾驶增加了伊朗高速公路上严重撞车的风险.
科学领域:
- 交通安全工程 交通安全工程
- 数据科学数据科学数据科学
- 运输研究 运输研究
背景情况:
- 机器学习 (ML) 模型越来越多地用于事故严重性预测.
- 这些ML模型的解释性经常被忽视,限制了实际见解.
- 了解导致碰撞严重性的因素对于有效的安全干预至关重要.
研究的目的:
- 实施可解释的机器学习技术,以可视化影响碰撞严重性的因素.
- 分析伊朗五年来的高速公路事故数据,以确定关键的风险因素.
- 通过清晰的模型解释,增强交通安全决策过程.
主要方法:
- 应用了各种机器学习模型:分类和回归树 (CART),K-最近邻居 (KNNs),随机森林 (RF),人工神经网络 (ANN) 和支持矢量机器 (SVM).
- 使用的累积局部效应 (ALE) 图表用于模型解释.
- 使用精度,回忆,F1得分和ROC指标评估模型性能,RF显示出优异的结果.
主要成果:
- 随机森林模型在预测撞车严重程度方面表现最高.
- 与严重碰撞相关的关键因素包括轻度交通条件 (临界值约为0.05或0.38).
- 大型卡车和公共汽车的比例较高 (例如,10%和4%),时速超过90公里,驾驶员年龄在30岁以下,翻车,碰撞固定物体/障碍物,夜间驾驶和驾驶员疲劳显著增加严重碰撞的可能性.
结论:
- 可解释的ML,特别是带有ALE的RF,有效地可视化和量化各种因素对碰撞严重性的影响.
- 调查结果为高速公路上的有针对性的安全措施提供了可操作的见解.
- 该研究强调了驾驶员人口统计学,车辆类型,环境条件和碰撞动态在确定严重性的重要性.
相关概念视频
Regression Analysis
5.7K
Regression analysis is a statistical tool that describes a mathematical relationship between a dependent variable and one or more independent variables.
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
5.7K
Determination of Expected Frequency
2.2K
Suppose one wants to test independence between the two variables of a contingency table. The values in the table constitute the observed frequencies of the dataset. But how does one determine the expected frequency of the dataset? One of the important assumptions is that the two variables are independent, which means the variables do not influence each other. For independent variables, the statistical probability of any event involving both variables is calculated by multiplying the individual...
2.2K
Relative Risk
151
Relative risk (RR) is a statistical measure commonly used in epidemiology to compare the likelihood of a particular event occurring between two groups. This metric is important for evaluating the relationship between exposure to a specific risk factor and the probability of a particular outcome. It plays a crucial role in medical research, public health studies, and risk assessment. Relative risk quantifies how much more (or less) likely an event is to occur in an exposed group compared to an...
151
Multiple Regression
3.0K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.0K
Hypothesis Test for Test of Independence
3.6K
The test of independence is a chi-square-based test used to determine whether two variables or factors are independent or dependent. This hypothesis test is used to examine the independence of the variables. One can construct two qualitative survey questions or experiments based on the variables in a contingency table. The goal is to see if the two variables are unrelated (independent) or related (dependent). The null and alternative hypotheses for this test are:
H0: The two variables (factors)...
H0: The two variables (factors)...
3.6K
Statistical Methods for Analyzing Epidemiological Data
361
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
361


