使用基于随机森林的Shapley添加式解释方法,跨不同事故严重程度的自行车撞车频率建模
Tao Li1, Ruiqi Wang1, Hongliang Ding2
1School of Transportation and Logistics, Southwest Jiaotong University, Chengdu, Sichuan, China.
International journal of injury control and safety promotion
|March 29, 2025
概括
了解自行车撞车风险因素是提高安全性的关键. 这项研究使用先进的建模来识别建筑密度和人口等关键因素,揭示它们如何影响碰撞严重程度和频率.
科学领域:
- 城市规划和交通安全.
- 统计建模和数据分析.
- 道路安全工程工程 道路安全工程
背景情况:
- 自行车事故研究往往缺乏对危险因素对碰撞严重性的影响的详细解释.
- 了解这些机制对于制定有效的安全干预措施至关重要.
- 现有的研究还没有充分探索各种因素在不同事故严重程度的相互作用.
研究的目的:
- 调查各种风险因素对不同严重程度的自行车撞车频率的影响.
- 应用先进的机器学习技术,以更深入地了解撞车原因.
- 确定和量化伦敦自行车碰撞严重性的主要决定因素.
主要方法:
- 利用了伦敦车祸数据 (2017-2019) 的三年时间.
- 采用随机森林和沙普利增量解释 (RF-SHAP) 用于预测建模和因素重要性分析.
- 关于人口人口统计,土地使用,道路基础设施和交通流动的综合数据.
主要成果:
- 确定了建筑面积比例和人口密度作为影响单车事故数量的关键因素.
- 量化了风险因素的主要和相互作用的影响,揭示了复杂的关系.
- 在道路网络连接2.25以下的交通流量和事故频率之间发现了负相关性.
- 在严重碰撞时确定了道路密度 (6.3) 的安全冲击边界.
- 观察到住宅区对轻伤事故的三级影响,控制人口密度.
结论:
- RF-SHAP有效地识别了关键的风险因素及其对自行车事故严重性的影响.
- 调查结果为城市环境中的有针对性的,具有成本效益的安全对策提供了关键的见解.
- 通过数据驱动的基础设施,交通和教育战略,可以实现对自行车安全的长期改进.
相关概念视频
Determination of Expected Frequency
2.1K
Suppose one wants to test independence between the two variables of a contingency table. The values in the table constitute the observed frequencies of the dataset. But how does one determine the expected frequency of the dataset? One of the important assumptions is that the two variables are independent, which means the variables do not influence each other. For independent variables, the statistical probability of any event involving both variables is calculated by multiplying the individual...
2.1K
Random Error
675
Random or indeterminate errors originate from various uncontrollable variables, such as variations in environmental conditions, instrument imperfections, or the inherent variability of the phenomena being measured. Usually, these errors cannot be predicted, estimated, or characterized because their direction and magnitude often vary in magnitude and direction even during consecutive measurements. As a result, they are difficult to eliminate. However, the aggregate effect of these errors can be...
675
Survival Tree
39
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
39
Introduction to Test of Independence
2.1K
In statistics, the term independence means that one can directly obtain the probability of any event involving both variables by multiplying their individual probabilities. Tests of independence are chi-square tests involving the use of a contingency table of observed (data) values.
The test statistic for a test of independence is similar to that of a goodness-of-fit test:
The test statistic for a test of independence is similar to that of a goodness-of-fit test:
2.1K
Parametric Survival Analysis: Weibull and Exponential Methods
271
Parametric survival analysis models survival data by assuming a specific probability distribution for the time until an event occurs. The Weibull and exponential distributions are two of the most commonly used methods in this context, due to their versatility and relatively straightforward application.
Weibull Distribution
The Weibull distribution is a flexible model used in parametric survival analysis. It can handle both increasing and decreasing hazard rates, depending on its shape parameter...
Weibull Distribution
The Weibull distribution is a flexible model used in parametric survival analysis. It can handle both increasing and decreasing hazard rates, depending on its shape parameter...
271
Random Variables
11.2K
A random variable is a single numerical value that indicates the outcome of a procedure. The concept of random variables is fundamental to the probability theory and was introduced by a Russian mathematician, Pafnuty Chebyshev, in the mid-nineteenth century.
Uppercase letters such as X or Y denote a random variable. Lowercase letters like x or y denote the value of a random variable. If X is a random variable, then X is written in words, and x is given as a number.
For example, let X = the...
Uppercase letters such as X or Y denote a random variable. Lowercase letters like x or y denote the value of a random variable. If X is a random variable, then X is written in words, and x is given as a number.
For example, let X = the...
11.2K


