使用可解释机器学习预测事故热点中的错误并调查时间,天气和行为因素:对远程学大数据的分析
Ali Golestani1, Nazila Rezaei1, Mohammad-Reza Malekpour1
1Non-Communicable Diseases Research Center, Endocrinology and Metabolism Population Sciences Institute, Tehran University of Medical Sciences, Tehran, Iran.
PloS one
|July 8, 2025
概括
这项研究使用可解释的机器学习来预测伊朗的道路事故热点,发现像省份和道路类型这样的空间因素是最关键的. 疲劳等行为错误也会增加事故风险,指导有针对性的安全干预.
科学领域:
- 交通安全研究 交通安全研究
- 公共卫生 公共卫生
- 数据科学在运输中的应用.
背景情况:
- 在全球范围内,道路交通事故 (RTA) 给公众健康和经济带来了重大挑战.
- 机器学习 (ML) 有助于RTA预测,但可解释性问题阻碍了政策实施.
- 本研究解决了在识别RTA热点及其原因时需要可解释的ML的需求.
研究的目的:
- 使用可解释的ML模型预测道路事故热点中的错误.
- 识别和解释伊朗RTA错误的关键预测因素.
- 为针对性干预和数据驱动的政策制定提供信息,以缓解RTA.
主要方法:
- 利用来自伊朗1673辆城际巴士 (2020) 的远程数据与天气数据相结合.
- 在619,988个记录上采用集体方法来训练6个ML模型 (逻辑回归,KNN,随机森林,XGBoost,naive Bayes,SVM).
- 选择XGBoost作为表现最好的模型,并使用SHAP来解释有影响力的预测因素.
主要成果:
- XGBoost获得了最高的性能,AUC为91.70%.
- 空间变量 (省份,道路类型) 是事故热点错误的最有影响力的预测因素.
- 行为因素 (疲劳) 和天气变量 (露点,湿度) 也很重要,而时间因素没有.
结论:
- 空间因素在很大程度上主导了道路交通事故热点错误的预测.
- 调查结果强调需要改善基础设施和数据驱动的政策,以减少RTA风险.
- 可解释的ML为有效的道路安全干预提供了宝贵的见解.
相关概念视频
Random Error
1.6K
Random or indeterminate errors originate from various uncontrollable variables, such as variations in environmental conditions, instrument imperfections, or the inherent variability of the phenomena being measured. Usually, these errors cannot be predicted, estimated, or characterized because their direction and magnitude often vary in magnitude and direction even during consecutive measurements. As a result, they are difficult to eliminate. However, the aggregate effect of these errors can be...
1.6K
Manipulation and Analysis
61
GIS manipulation and analysis functions are vital for decision-making and planning. These activities range from data retrieval tasks, such as selecting information based on specific criteria, to advanced analytical techniques that address complex spatial problems.One critical GIS analysis method is overlaying, which combines multiple data layers to examine impacts. For example, overlaying a river-dammed lake boundary with road networks can identify affected infrastructure. Another common...
61
Steps in Outbreak Investigation
209
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
209
Errors in Global Positioning System
115
Global Positioning System (GPS) technology has revolutionized navigation and positioning, but its accuracy is often compromised by various errors. These errors, stemming from environmental, satellite, and receiver-related factors, require careful mitigation to ensure reliable performance across applications.Atmospheric ErrorsGPS signals travel through the Earth’s ionosphere and troposphere, introducing delays which affect accuracy. The ionosphere is strongly influenced by charged particles,...
115
Hypothesis Test for Test of Independence
3.7K
The test of independence is a chi-square-based test used to determine whether two variables or factors are independent or dependent. This hypothesis test is used to examine the independence of the variables. One can construct two qualitative survey questions or experiments based on the variables in a contingency table. The goal is to see if the two variables are unrelated (independent) or related (dependent). The null and alternative hypotheses for this test are:
H0: The two variables (factors)...
H0: The two variables (factors)...
3.7K
Determination of Expected Frequency
2.3K
Suppose one wants to test independence between the two variables of a contingency table. The values in the table constitute the observed frequencies of the dataset. But how does one determine the expected frequency of the dataset? One of the important assumptions is that the two variables are independent, which means the variables do not influence each other. For independent variables, the statistical probability of any event involving both variables is calculated by multiplying the individual...
2.3K


