Related Experiment Video
Updated: May 9, 2026

11:41
Evaluation of an Exclusive Spur Dike U-Turn Design with Radar-Collected Data and Simulation
Published on: February 1, 2020
20.3K
A large language model framework to uncover underreporting in traffic crashes.
Cristian Arteaga1, JeeWoong Park1
1Department of Civil and Environmental Engineering, University of Nevada Las Vegas, USA.
Journal of Safety Research
|February 22, 2025
Summary
This study introduces a framework using Large Language Models (LLMs) to automatically identify underreported factors in traffic crash data, improving safety analysis efficiency and accuracy.
Area of Science:
- Traffic Safety
- Data Science
- Natural Language Processing
Background:
- Traffic crash reports are vital for developing safety countermeasures.
- Underreporting of crash factors due to data collection errors is a significant issue.
- Manual data correction is time-consuming and prone to errors, especially for large datasets.
Purpose of the Study:
- To develop and evaluate a framework for analyzing traffic crash narratives.
- To uncover underreported crash factors using Large Language Models (LLMs).
- To improve the efficiency and accuracy of traffic safety data analysis.
Main Methods:
- The framework integrates prompt engineering, LLM parameter selection, output parsing, and underreporting determination.
- A case study focused on identifying underreported alcohol involvement in traffic crashes.
- Evaluated performance across different LLMs (ChatGPT, Flan-UL2, Llama-2), prompt types, and generation parameters using 500 Massachusetts crash reports.
Main Results:
- The framework achieved high recall (up to 1.0) and precision (up to 0.93) in identifying underreported crash instances.
- Demonstrated efficient and accurate uncovering of underreporting in crash data.
- The approach does not require extensive natural language processing expertise from safety analysts.
Conclusions:
- The developed framework effectively addresses critical gaps in traffic safety analysis.
- Offers a novel method to enhance the quality and comprehensiveness of traffic crash records.
- Paves the way for more effective traffic safety countermeasure development.
Related Concept Videos
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
Models, Theories, and Laws
Scientists frequently use models to help them comprehend a specific collection of phenomena. In physics, a model is a condensed version of a physical system that is too complex to study thoroughly. One such example is the light wave model; unlike water waves, light waves are typically invisible to us. Nonetheless, it is helpful to think of light as being composed of waves, since investigations show that light behaves like water waves. Since it is impossible to visually see what is genuinely...

