Related Experiment Video
Updated: Mar 15, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
DRPVLM: A generative multimodal large language model for real-time driving risk prediction
Junhua Wang1, Wenhao Zhang1, Ting Fu1
1The Key Laboratory of Road and Traffic Engineering, Ministry of Education, Tongji University, Shanghai 201804, China; College of Transportation, Tongji University, 4800 Cao'an Highway, Shanghai 201804, China.
Large language models (LLMs) significantly improve driving risk prediction by analyzing visual and sensor data. This technology enhances driver safety and understanding of traffic environments.
Area of Science:
- Artificial Intelligence
- Computer Vision
- Transportation Safety
Background:
- Large language models (LLMs) possess advanced comprehension abilities but their application in enhancing driver understanding of traffic environments and identifying driving risks is underexplored.
- Current in-vehicle systems have limited capabilities in real-time driving risk assessment.
Purpose of the Study:
- To propose and evaluate a Driving Risk Prediction Vision-Language Model (DRPVLM) for recognizing real-time driving risks.
- To assess the performance enhancement offered by multimodal LLMs in driving risk prediction compared to traditional sensor-based methods.
Main Methods:
- Fine-tuned several open-source multimodal LLMs (Qwen-2.5-VL, Gemma-3, Llama-3.2-Vision) using LoRA.
- Processed video and image data from the Shanghai Naturalistic Driving Study to extract multi-dimensional features (road environment, traffic conditions, driver states).
- Integrated LLM-extracted features with structured trajectory data into a Long Short-Term Memory (LSTM) network for risk prediction.
Main Results:
- Multimodal LLMs significantly improved driving risk prediction accuracy, with Qwen2.5-VL-32B achieving 0.89-0.92 accuracy and 0.88-0.91 F1 score.
- All tested LLMs outperformed the baseline model relying solely on structured trajectory data, which dropped below 0.7 accuracy for longer prediction horizons.
- Feature importance analysis confirmed the meaningful contribution of LLM-extracted variables in supplementing trajectory data.
Conclusions:
- Multimodal LLMs are effective in enhancing feature extraction for driving risk prediction.
- The proposed DRPVLM framework demonstrates strong potential for real-time driving risk prediction, improving overall road safety.
Related Concept Videos
Multi-input and Multi-variable systems
In the absence of...
Language and Cognition
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Improving Translational Accuracy
Improving Translational Accuracy
