Related Experiment Video
Updated: Jun 17, 2026

A Prediction Error-driven Retrieval Procedure for Destabilizing and Rewriting Maladaptive Reward Memories in Hazardous Drinkers
Published on: January 5, 2018
Improving crash data quality by identifying misclassified alcohol-involved crashes using NLP on narrative data
Sudesh Bhagat1, Raghupathi Kandiboina2, Ibne Farabi Shihab3
1Department of Civil Construction and Environmental Engineering, Iowa State University of Science and Technology, Ames, IA 50011-1066, USA.
Introduction:
Road traffic crashes remain a leading cause of fatalities worldwide, underscoring the need for accurate data to guide prevention strategies and evidence-based policymaking. However, crash databases often suffer from misclassification, underreporting, and inconsistencies, particularly in alcohol-involved cases, which limits the reliability of safety analyses.
Method:
This study addresses this issue by identifying and quantifying Misclassified Alcohol-Involved Crashes (MAICs) using a Natural Language Processing (NLP) framework based on the BERT model. The framework analyzed 371,062 crash records from Iowa (2016-2022) and identified 3,895 misclassified alcohol-involved crashes (MAICs) out of 19,177 alcohol-involved cases predicted by the model, corresponding to an overall misclassification rate of 20.35% and a confidence interval of 18.86%-21.85%. To examine the factors contributing to these errors, a mixed-effects Probit Logit regression model was applied, incorporating behavioral, environmental, and roadway attributes.
Results:
Results indicated that fatal and nighttime crashes were less likely to be misclassified, whereas crashes involving older or younger drivers, heavy trucks, and vulnerable road users showed higher odds of misclassification. A Local Indicators of Spatial Association (LISA) analysis revealed significant county-level clusters of misclassifications, suggesting regional differences in enforcement and reporting practices.
Related Concept Videos
Hypothesis Test for Test of Independence
H0: The two variables (factors)...
Improving Translational Accuracy
Improving Translational Accuracy
Alcohols from Carbonyl Compounds: Reduction
Catalytic hydrogenation is similar to the reduction of an alkene or alkyne by adding H2 across the pi bond in the presence of transition metal catalysts like Raney Ni, Pd–C, Pt, or Ru. Aldehydes and ketones can be reduced by this method, often under mild to moderate heat (25–100°C) and...
Types of Collisions - II
How Data are Classified: Categorical Data
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...