Related Experiment Video
Updated: Jul 9, 2025

11:41
Evaluation of an Exclusive Spur Dike U-Turn Design with Radar-Collected Data and Simulation
Published on: February 1, 2020
20.4K
Advancing proactive crash prediction: A discretized duration approach for predicting crashes and severity
Diwas Thapa1, Sabyasachee Mishra1, Nagendra R Velaga2
1Department of Civil Engineering, University of Memphis, Memphis, TN 38152, United States.
Accident; Analysis and Prevention
|December 6, 2023
Summary
This study introduces a novel duration-based framework for predicting traffic crashes and their severity. A 15% data sample offers a balance between accuracy and efficiency in crash prediction modeling.
Area of Science:
- Transportation Engineering
- Statistical Modeling
- Road Safety
Background:
- Proactive crash prediction models increasingly use machine learning and AI.
- Statistical models offer causal insights and effect size estimation.
- Existing statistical models often use a case-control approach with challenges in ratio definition and severity incorporation.
Purpose of the Study:
- To extend the duration-based modeling framework for predicting traffic crashes and their severity.
- To investigate the trade-off between model performance and estimation time when incorporating crash severities.
- To identify optimal data sampling strategies for robust crash prediction.
Main Methods:
- Developed a novel duration-based framework integrating crash severity.
- Explored data sampling strategies (15% epoch-level sample) to manage computational complexity.
- Conducted stability analysis of predictor variables across different sample sizes.
Main Results:
- A 15% epoch-level sample provides a balanced approach between data size and predictive accuracy.
- Certain variables (e.g., Time of day, Weather, Lighting, Volume) require larger samples for stable estimation.
- Other variables (e.g., Daytime, Terrain, Land use, Number of lanes, Speed) converge with smaller sample increases.
- The model shows better performance on highway segments with shorter crash durations (<100 hours).
Conclusions:
- The proposed duration-based framework effectively predicts crashes and severity.
- Epoch-level data sampling, specifically 15%, is a viable strategy for balancing model performance and computational efficiency.
- Understanding variable sensitivity to sample size aids in optimizing data collection and model development for road safety.
Related Concept Videos
Determination of Expected Frequency
2.2K
Suppose one wants to test independence between the two variables of a contingency table. The values in the table constitute the observed frequencies of the dataset. But how does one determine the expected frequency of the dataset? One of the important assumptions is that the two variables are independent, which means the variables do not influence each other. For independent variables, the statistical probability of any event involving both variables is calculated by multiplying the individual...
2.2K
Hazard Rate
114
The hazard rate, also known as the hazard function or failure rate, is a statistical measure used to describe the instantaneous rate at which an event occurs, given that the event has not yet happened. From a probabilistic perspective, it represents the likelihood that a subject will experience the event in a very small time interval, conditional on surviving up to the beginning of that interval. In terms of frequency, the hazard rate can be viewed as the ratio of the number of events to the...
114
Censoring Survival Data
100
Survival analysis is a statistical method used to analyze time-to-event data, often employed in fields such as medicine, engineering, and social sciences. One of the key challenges in survival analysis is dealing with incomplete data, a phenomenon known as "censoring." Censoring occurs when the event of interest (such as death, relapse, or system failure) has not occurred for some individuals by the end of the study period or is otherwise unobservable, and it might have many different...
100
Survival Tree
87
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
87
Prediction Intervals
2.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.3K
Introduction To Survival Analysis
247
Survival analysis is a statistical method used to study time-to-event data, where the "event" might represent outcomes like death, disease relapse, system failure, or recovery. A unique feature of survival data is censoring, which occurs when the event of interest has not been observed for some individuals during the study period. This requires specialized techniques to handle incomplete data effectively.
The primary goal of survival analysis is to estimate survival time—the time...
The primary goal of survival analysis is to estimate survival time—the time...
247

