Related Experiment Videos
Benchmarking Zero-Shot Open-Vocabulary and Fine-Tuned Object Detectors for Underground Mine Personnel Detection
Ellen Essien1, Samuel Frimpong1
1Department of Mining and Explosives Engineering, Missouri University of Science and Technology, Rolla, MO 65409, USA.
Abstract:
Reliable personnel detection is critical for the safe deployment of autonomous haulage systems in underground mining, where challenging environmental conditions demand robust real-time perception. Existing research has focused primarily on fine-tuned convolutional detectors, while systematic comparisons with zero-shot vision-language models remain limited. This study presents a cross-paradigm benchmark comparing four zero-shot vision-language models (YOLO-World, Grounding DINO, OWL-ViT, and OWLv2) with four fine-tuned YOLO detectors (YOLOv8s, YOLOv9s, YOLO11s, and YOLO26s) using 31,396 real-world underground coal mine images. Detection performance was evaluated using precision, recall, F1-score, average precision, inference speed, and condition- and target scale-specific recall. The experimental results show that the fine-tuned detectors achieved F1-scores of 0.8449-0.8662 and AP50 values of 0.8791-0.9135, with YOLO26s achieving the strongest overall performance. In comparison, the zero-shot models achieved F1-scores of 0.3186-0.5484 and AP50 values of 0.2484-0.5108, with Grounding DINO performing best among the zero-shot models. Fine-tuned detectors also maintained substantially higher recall for occluded personnel and small apparent targets. These findings demonstrate a substantial performance advantage for domain-specific fine-tuning over the evaluated zero-shot approaches and establish a controlled cross-paradigm benchmark for comparing detection paradigms for underground personnel perception.