Related Experiment Video
Updated: Mar 15, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
DRPVLM: A generative multimodal large language model for real-time driving risk prediction
Junhua Wang1, Wenhao Zhang1, Ting Fu1
1The Key Laboratory of Road and Traffic Engineering, Ministry of Education, Tongji University, Shanghai 201804, China; College of Transportation, Tongji University, 4800 Cao'an Highway, Shanghai 201804, China.
Abstract:
Large language models (LLMs), known for their general knowledge comprehension capabilities, have recently been integrated into certain in-vehicle systems. However, their potential to enhance driver understanding of traffic environments and support driving risk identification remains underexplored. This study proposes a Driving Risk Prediction Vision-Language Model (DRPVLM) to recognize real-time driving risks. The framework is fine-tuned using LoRA on several open-source multimodal LLMs of different parameter sizes, including three Qwen-2.5-VL models (32B, 7B, and 3B), Gemma-3-12B-it, and Llama-3.2-11B-Vision. DRPVLM processes video and image data from the Shanghai Naturalistic Driving Study to extract multi-dimensional features, including road environment, traffic conditions, and driver states, which complement structured trajectory data obtained from in-vehicle sensors. These features are subsequently fed into a Long Short-Term Memory (LSTM) neural network for risk prediction. In addition, we compare DRPVLM, equipped with each of these multimodal LLMs, with the model using only structured trajectory data collected from in-vehicle sensors to evaluate their predictive performance. Results indicate that Multimodal LLMs significantly enhance driving-risk prediction, with fine-tuned Qwen2.5-VL-32B achieving an accuracy of 0.89 to 0.92 and an F1 score of 0.88 to 0.91 across observation windows. LLMs with different parameter sizes also perform well and clearly outperforming the baseline model that relies solely on structured trajectory data collected from in-vehicle sensors, whose performance drops below 0.7 for longer prediction horizons. Feature importance analysis shows that all five LLM‑extracted variables make meaningful contributions, effectively supplementing structured trajectory features. These findings demonstrate the effectiveness of multimodal LLMs in enhancing risk feature extraction and improving driving risk prediction performance, highlighting the strong potential of LLMs for real-time driving risk prediction.
Related Concept Videos
Multi-input and Multi-variable systems
In the absence of...
Language and Cognition
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Improving Translational Accuracy
Improving Translational Accuracy
