Related Experiment Video
Updated: Jan 26, 2026

Evaluating the Effect of Roadside Parking on a Dual-Direction Urban Street
Published on: January 20, 2023
From crash reports to safer roads: a multimodal framework integrating vision-language models and street view analysis
Guanhe Wu1, Xintong Yan1, Yajie Zou2
1Department of Civil and Environmental Engineering, University of Massachusetts Lowell, Lowell, MA 01854, USA.
None:
Crash narratives and diagrams contain rich causal and contextual information that is often underutilized in traditional road safety analysis, which typically relies on aggregated crash counts and tabulated variables. Such approaches have limited ability to explain the mechanisms of individual crashes. This study proposes a mechanism-oriented, multimodal framework for scalable and interpretable roadway risk diagnosis. The framework employs a vision-language model to jointly analyze each crash narrative and diagram, constructing an explicit reasoning chain that distinguishes immediate crash-triggering behaviors from underlying contextual constraints. Based on inferred root cause keywords, each crash is decomposed into proportional contributions of five predefined factors: human, vehicle, road, environment, and traffic signal. These attributions are spatially aggregated to identify hotspots with elevated road-related risk. For high-risk locations, street-view imagery is analyzed in the crash context to diagnose potential roadway design deficiencies and generate targeted recommendations. The framework is evaluated using 4,302 crash reports from Massachusetts with expert-annotated attributions. Vision-language models substantially outperform locally trained regression baselines, particularly under imbalanced factor distributions. Among all tested models, the Grok 2 multimodal model achieves the highest agreement with expert annotations. Further analysis shows that the distribution of inferred root cause keywords in the factor attribution space effectively explains how relative factor contributions are formed. The street-view-based diagnostic outputs align closely with real-world engineering issues, demonstrating that combining causal reasoning with multimodal crash data enables practical, mechanism-driven road safety diagnostics beyond frequency-based screening.
Related Concept Videos
Vision
Language
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...
Design Example: Alignment of a Road Line Using GIS
Color Vision
Components of Language
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...

