Resource-efficient fine-tuning of large vision-language models for multimodal perception in autonomous excavators

Hung Viet Nguyen1, Hyojin Park2, Namhyun Yoo3

  • 1Department of Digital Anti-aging Healthcare, INJE University, Kimhae, Republic of Korea.

PubMed
Summary

This study introduces an efficient method for fine-tuning large vision-language models (LVLMs) for autonomous excavators. The approach enhances human/obstacle detection and weather classification on standard hardware.