Related Experiment Video
Updated: Jun 17, 2026

A Prediction Error-driven Retrieval Procedure for Destabilizing and Rewriting Maladaptive Reward Memories in Hazardous Drinkers
Published on: January 5, 2018
Improving crash data quality by identifying misclassified alcohol-involved crashes using NLP on narrative data
Sudesh Bhagat1, Raghupathi Kandiboina2, Ibne Farabi Shihab3
1Department of Civil Construction and Environmental Engineering, Iowa State University of Science and Technology, Ames, IA 50011-1066, USA.
Accurate road safety data is crucial. This study used Natural Language Processing (NLP) to identify misclassified alcohol-involved crashes, finding a 20.35% misclassification rate and key contributing factors.
Area of Science:
- Road safety research
- Transportation engineering
- Data science
Background:
- Road traffic crashes are a major global cause of death.
- Inaccurate crash data, especially for alcohol-involved incidents, hinders effective safety strategies.
- Existing databases face challenges like misclassification and underreporting.
Purpose of the Study:
- To identify and quantify misclassified alcohol-involved crashes (MAICs).
- To analyze factors contributing to misclassification errors in crash data.
- To improve the reliability of road safety analyses.
Main Methods:
- Utilized a Natural Language Processing (NLP) framework based on the BERT model.
- Analyzed 371,062 crash records from Iowa (2016-2022).
- Employed mixed-effects Probit Logit regression and Local Indicators of Spatial Association (LISA) analyses.
Main Results:
- Identified a 20.35% misclassification rate for alcohol-involved crashes.
- Fatal and nighttime crashes were less likely to be misclassified.
- Crashes involving specific driver groups (older/younger), heavy trucks, and vulnerable road users had higher misclassification odds.
- Discovered significant county-level clusters of misclassifications, indicating regional variations.
Conclusions:
- Natural Language Processing (NLP) effectively identifies misclassified alcohol-involved crashes.
- Driver demographics, vehicle types, and crash timing influence misclassification.
- Spatial analysis highlights regional disparities in crash reporting and enforcement practices.
Related Concept Videos
Hypothesis Test for Test of Independence
H0: The two variables (factors)...
Improving Translational Accuracy
Improving Translational Accuracy
Alcohols from Carbonyl Compounds: Reduction
Catalytic hydrogenation is similar to the reduction of an alkene or alkyne by adding H2 across the pi bond in the presence of transition metal catalysts like Raney Ni, Pd–C, Pt, or Ru. Aldehydes and ketones can be reduced by this method, often under mild to moderate heat (25–100°C) and...
Types of Collisions - II
How Data are Classified: Categorical Data
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...