分析影响铁路事故的因素:使用多项物流回归和数据挖掘的预测方法
Jaroslav Mašek1, Lucia Duricova2, Juraj Čamaj1
1Department of Railway Transport, Faculty of Operation and Economics of Transport and Communications, University of Zilina, Zilina, Slovak Republic.
PloS one
|October 7, 2025
概括
铁路自杀是最常见的铁路事故,与社会经济因素如利息,婚姻和生育率有关. 这项研究使用数据挖掘预测事故率,为铁路安全和自杀预防提供了洞察力.
科学领域:
- 社会学 社会学 社会学
- 运输安全运输安全
- 数据科学数据科学数据科学
背景情况:
- 铁路事故,特别是自杀事件,造成了严重的运营中断和基础设施损坏.
- 与自杀有关的事件是最常见的铁路事故类型.
- 了解社会经济因素和铁路自杀之间的联系对于安全至关重要.
研究的目的:
- 确定与自杀有关的铁路事件与社会经济因素之间的关系.
- 基于社会经济数据,开发铁路事故率的预测模型.
- 为改善铁路安全和预防自杀策略提供见解.
主要方法:
- 使用斯洛伐克共和国铁路 (2015-2022) 的数据.
- 使用数据挖掘方法来分析事件模式.
- 开发了一个后勤回归模型来预测事故率.
主要成果:
- 确定了关键的社会经济预测因素:利率,婚姻率和生育率.
- 后勤回归模型展示了高预测性能.
- 数据挖掘方法有效地提取了相关的模式和关系.
结论:
- 社会影响,反映在社会经济因素中,在与铁路有关的自杀事件中发挥着关键作用.
- 该预测模型为加强铁路安全协议提供了实际指导.
- 研究结果支持铁路环境中的有针对性的自杀预防工作.
相关概念视频
Multiple Regression
3.8K
Multiple regression assesses a linear relationship between one response or dependent variable and two or more independent variables. It has many practical applications.
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
3.8K
Hazard Rate
401
The hazard rate, also known as the hazard function or failure rate, is a statistical measure used to describe the instantaneous rate at which an event occurs, given that the event has not yet happened. From a probabilistic perspective, it represents the likelihood that a subject will experience the event in a very small time interval, conditional on surviving up to the beginning of that interval. In terms of frequency, the hazard rate can be viewed as the ratio of the number of events to the...
401
Determination of Expected Frequency
2.5K
Suppose one wants to test independence between the two variables of a contingency table. The values in the table constitute the observed frequencies of the dataset. But how does one determine the expected frequency of the dataset? One of the important assumptions is that the two variables are independent, which means the variables do not influence each other. For independent variables, the statistical probability of any event involving both variables is calculated by multiplying the individual...
2.5K
Introduction to Test of Independence
2.9K
In statistics, the term independence means that one can directly obtain the probability of any event involving both variables by multiplying their individual probabilities. Tests of independence are chi-square tests involving the use of a contingency table of observed (data) values.
The test statistic for a test of independence is similar to that of a goodness-of-fit test:
The test statistic for a test of independence is similar to that of a goodness-of-fit test:
2.9K
Hypothesis Test for Test of Independence
7.4K
The test of independence is a chi-square-based test used to determine whether two variables or factors are independent or dependent. This hypothesis test is used to examine the independence of the variables. One can construct two qualitative survey questions or experiments based on the variables in a contingency table. The goal is to see if the two variables are unrelated (independent) or related (dependent). The null and alternative hypotheses for this test are:
H0: The two variables (factors)...
H0: The two variables (factors)...
7.4K
Survival Tree
385
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
385

