Related Experiment Video
Updated: Jan 9, 2026

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
Analyzing crash severity through impairment and protection: A hybrid XGBoost-Bayesian network approach
Ashutosh Dumka1, Raghupathi Kandiboina1, Aparna Joshi1
1Department of Civil Construction and Environmental Engineering, Iowa State University of Science and Technology, Ames, Iowa.
Objective:
This study analyzes how impairment, protection status, vehicle type, age group, and road characteristics collectively influence crash severity. While previous research has examined these factors in isolation, this study adopts a hybrid framework combining XGBoost to identify key predictors with Bayesian networks for modeling conditional dependencies and estimating risk under various scenarios. Additionally, the study introduces a temporal dimension by comparing crash severity patterns across pre-COVID and post-COVID periods using relative risk scores. This integrated approach supports data-driven, scenario-specific, and time-sensitive safety interventions to mitigate crash severity and improve road safety.
Methods:
Historical crash data from Iowa (26,111 records for fatal, major, and minor crashes) were analyzed using 14 variables covering infrastructure, environment, driver, and vehicle characteristics. An XGBoost model was applied for feature selection, with SHAP values used to interpret the most influential predictors. A Bayesian network, built using the tree-augmented naive Bayes method, was then used for probabilistic inference. The Bayesian network's conditional probability tables were leveraged to compute relative risk scores under various evidence settings. A temporal analysis was conducted by segmenting the data into pre-COVID and post-COVID periods, enabling a comparative assessment of evolving crash risk patterns across user categories.
Results:
The study identifies impairment and lack of protection as key drivers of crash severity, with relative risk scores differentiating high-risk groups. The temporal analysis comparing pre- and post-COVID periods reveals a consistent rise in relative risk scores post pandemic, underscoring shifts in crash risk patterns and reinforcing the need for adaptive, data-driven safety interventions over time.
Conclusions:
This study examines the interplay among protection status, impairment, driver demographics, vehicle type, and road characteristics, providing a deeper understanding of how these factors collectively influence crash severity. In addition to analyzing these complex relationships, the study introduces a probabilistic and robust framework that moves beyond traditional regression models. By integrating XGBoost for feature selection and Bayesian networks for conditional probability modeling, this approach captures conditional dependencies among key factors, enhancing interpretability. In addition to scenario-based relative risk estimation, the study introduces a temporal component by comparing pre-COVID and post-COVID crash patterns. The findings support data-driven, targeted interventions and policies, offering a flexible and interpretable tool for monitoring crash severity trends and guiding effective safety strategies over time.
Related Concept Videos
Design Example: Analyzing Capacity Contours for Flood Risk Assessment
Survival Tree
Building a Survival Tree
Constructing a...
Quantifying and Rejecting Outliers: The Grubbs Test
Hazard Ratio
For example, in a clinical trial...
Parametric Survival Analysis: Weibull and Exponential Methods
Weibull Distribution
The Weibull distribution is a flexible model used in parametric survival analysis. It can handle both increasing and decreasing hazard rates, depending on its shape parameter...
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
