Improving crash data quality by identifying misclassified alcohol-involved crashes using NLP on narrative data

Sudesh Bhagat1, Raghupathi Kandiboina2, Ibne Farabi Shihab3

  • 1Department of Civil Construction and Environmental Engineering, Iowa State University of Science and Technology, Ames, IA 50011-1066, USA.

Summary

Accurate road safety data is crucial. This study used Natural Language Processing (NLP) to identify misclassified alcohol-involved crashes, finding a 20.35% misclassification rate and key contributing factors.

Related Concept Videos

Hypothesis Test for Test of Independence01:16

Hypothesis Test for Test of Independence

The test of independence is a chi-square-based test used to determine whether two variables or factors are independent or dependent. This hypothesis test is used to examine the independence of the variables. One can construct two qualitative survey questions or experiments based on the variables in a contingency table. The goal is to see if the two variables are unrelated (independent) or related (dependent). The null and alternative hypotheses for this test are:
H0: The two variables (factors)...
Improving Translational Accuracy02:07

Improving Translational Accuracy

Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
Improving Translational Accuracy02:07

Improving Translational Accuracy

Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
Alcohols from Carbonyl Compounds: Reduction02:23

Alcohols from Carbonyl Compounds: Reduction

Reduction is a simple strategy to convert a carbonyl group to a hydroxyl group. The three major pathways to reduce carbonyls to alcohols are catalytic hydrogenation, hydride reduction, and borane reduction.
Catalytic hydrogenation is similar to the reduction of an alkene or alkyne by adding H2 across the pi bond in the presence of transition metal catalysts like Raney Ni, Pd–C, Pt, or Ru. Aldehydes and ketones can be reduced by this method, often under mild to moderate heat (25–100°C) and...
Types of Collisions - II01:19

Types of Collisions - II

When two or more objects collide with each other, they can stick together to form one single composite object (after collision). The total mass of the object after the collision is the sum of the masses of the original objects, and it moves with a velocity dictated by the conservation of momentum. Although the system's total momentum remains constant, the kinetic energy decreases, and thus such a collision is an inelastic collision. Most of the collisions between objects in daily life are...
How Data are Classified: Categorical Data01:11

How Data are Classified: Categorical Data

A variable, usually notated by capital letters such as X and Y, is a characteristic or measurement that can be determined for each member of a population. Data are the actual values of variables. They may be numbers, or they may be words. Datum is a single value.
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...