Related Experiment Video
Updated: Jul 21, 2026

08:25
Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
Published on: May 7, 2019
9.5K
Resource-efficient fine-tuning of large vision-language models for multimodal perception in autonomous excavators
Hung Viet Nguyen1, Hyojin Park2, Namhyun Yoo3
1Department of Digital Anti-aging Healthcare, INJE University, Kimhae, Republic of Korea.
Frontiers in Artificial Intelligence
|December 4, 2025
Summary
This study introduces an efficient method for fine-tuning large vision-language models (LVLMs) for autonomous excavators. The approach enhances human/obstacle detection and weather classification on standard hardware.
Area of Science:
- Computer Vision
- Artificial Intelligence
- Robotics
Background:
- Large Vision-Language Models (LVLMs) offer advanced multimodal understanding but are underexplored in construction.
- Real-world autonomous excavator operations require robust visual recognition for safety and efficiency.
Purpose of the Study:
- To develop a resource-efficient framework for fine-tuning LVLMs for autonomous excavator tasks.
- To enable robust detection of humans and obstacles and accurate weather classification on consumer-grade hardware.
Main Methods:
- Leveraged Quantized Low-Rank Adaptation (QLoRA) and the Unsloth framework for efficient LVLM fine-tuning.
- Evaluated five open-source LVLMs (Llama-3.2-Vision, Qwen2-VL, Qwen2.5-VL, LLaVA-1.6, Gemma 3) on a domain-specific excavator-vision dataset.
- Fine-tuned models on 1,000 annotated frames and tested on 2,000 images.
Main Results:
- Achieved significant improvements in object detection and weather classification across evaluated LVLMs.
- Qwen2-VL-7B demonstrated superior performance with mAP@50 of 88.03% and accuracy of 84.54%.
- The fine-tuned Qwen2-VL-7B model robustly detects humans/obstacles and classifies weather.
Conclusions:
- The proposed framework enables efficient LVLM fine-tuning for autonomous excavator operations on consumer hardware.
- LVLM-based multimodal AI agents are feasible for enhancing safety, monitoring, and planning in construction environments.
- This research paves the way for advanced AI applications in heavy machinery automation.
