Related Experiment Video
Updated: Jul 5, 2025

04:57
Establishing a Competing Risk Regression Nomogram Model for Survival Data
Published on: October 23, 2020
10.2K
Fatal crashes and rare events logistic regression: an exploratory empirical study
Yuxie Xiao1,2, Lulu Lin1, Hanchu Zhou3
1School of Public Health, Sun Yat-sen University, Guangzhou, China.
Frontiers in Public Health
|January 22, 2024
Summary
Fatal road accidents are rare, making them hard to predict. A rare events logistic model (RELM) offers more accurate fatal crash predictions than the classic logit model (LM).
Area of Science:
- Traffic Safety Analytics
- Statistical Modeling
- Transportation Engineering
Background:
- Fatal road accidents are infrequent events, challenging traditional statistical models.
- Accurate prediction of fatal crashes is crucial for effective traffic safety interventions.
Purpose of the Study:
- To evaluate the effectiveness of a rare events logistic model (RELM) for predicting fatal road accidents.
- To compare the predictive accuracy of RELM against the classic logit model (LM).
Main Methods:
- Both logit model (LM) and rare events logistic model (RELM) were applied to analyze fatal crash data.
- Crash-injury datasets from Hillsborough County, Florida, were utilized for empirical evaluation.
- Model performance was assessed using Receiver Operating Characteristic (ROC) curves and Area Under the Curve (AUC) metrics.
Main Results:
- The rare events logistic model (RELM) demonstrated superior accuracy in predicting fatal crashes compared to the classic logit model (LM).
- Empirical analysis, supported by higher AUC values, indicated a clear advantage for RELM in predictive performance.
Conclusions:
- The rare events logistic model (RELM) is a more proficient tool for predicting fatal crashes than the classic logit model (LM).
- RELM is recommended for advanced traffic safety analytics requiring precise estimation of rare events.
Related Concept Videos
Determination of Expected Frequency
2.2K
Suppose one wants to test independence between the two variables of a contingency table. The values in the table constitute the observed frequencies of the dataset. But how does one determine the expected frequency of the dataset? One of the important assumptions is that the two variables are independent, which means the variables do not influence each other. For independent variables, the statistical probability of any event involving both variables is calculated by multiplying the individual...
2.2K
Hypothesis Test for Test of Independence
3.6K
The test of independence is a chi-square-based test used to determine whether two variables or factors are independent or dependent. This hypothesis test is used to examine the independence of the variables. One can construct two qualitative survey questions or experiments based on the variables in a contingency table. The goal is to see if the two variables are unrelated (independent) or related (dependent). The null and alternative hypotheses for this test are:
H0: The two variables (factors)...
H0: The two variables (factors)...
3.6K
Introduction to Test of Independence
2.3K
In statistics, the term independence means that one can directly obtain the probability of any event involving both variables by multiplying their individual probabilities. Tests of independence are chi-square tests involving the use of a contingency table of observed (data) values.
The test statistic for a test of independence is similar to that of a goodness-of-fit test:
The test statistic for a test of independence is similar to that of a goodness-of-fit test:
2.3K
Censoring Survival Data
96
Survival analysis is a statistical method used to analyze time-to-event data, often employed in fields such as medicine, engineering, and social sciences. One of the key challenges in survival analysis is dealing with incomplete data, a phenomenon known as "censoring." Censoring occurs when the event of interest (such as death, relapse, or system failure) has not occurred for some individuals by the end of the study period or is otherwise unobservable, and it might have many different...
96
Introduction To Survival Analysis
239
Survival analysis is a statistical method used to study time-to-event data, where the "event" might represent outcomes like death, disease relapse, system failure, or recovery. A unique feature of survival data is censoring, which occurs when the event of interest has not been observed for some individuals during the study period. This requires specialized techniques to handle incomplete data effectively.
The primary goal of survival analysis is to estimate survival time—the time...
The primary goal of survival analysis is to estimate survival time—the time...
239
Hazard Rate
112
The hazard rate, also known as the hazard function or failure rate, is a statistical measure used to describe the instantaneous rate at which an event occurs, given that the event has not yet happened. From a probabilistic perspective, it represents the likelihood that a subject will experience the event in a very small time interval, conditional on surviving up to the beginning of that interval. In terms of frequency, the hazard rate can be viewed as the ratio of the number of events to the...
112

