Related Experiment Video
Updated: May 5, 2026

Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
Published on: May 7, 2019
A Vision Language-Based Framework for Detecting Industrial Mechanical, Electrical, and Plumbing Assets Using
Masoud Kamali1, Behnam Atazadeh1, Abbas Rajabifard1
1The Centre for Spatial Data Infrastructures and Land Administration, Department of Infrastructure Engineering, The University of Melbourne, Melbourne, VIC 3010, Australia.
This study introduces a novel method for detecting unseen mechanical, electrical, and plumbing (MEP) assets using vision language models and object detectors. The approach significantly improves open-vocabulary detection of complex industrial assets.
Area of Science:
- Computer Vision
- Artificial Intelligence
- Industrial Asset Management
Background:
- Object detection models struggle with limited training data diversity and generalizing to new asset categories in industrial settings.
- Mechanical, electrical, and plumbing (MEP) assets present unique spatial and geometric complexities challenging current detection methods.
Purpose of the Study:
- To develop an effective approach for detecting unseen MEP assets in industrial environments.
- To leverage pre-trained vision language models and close-set object detectors for open-vocabulary asset detection using unlabeled data.
Main Methods:
- Utilized pre-trained vision language models and close-set object detectors.
- Employed Grounding DINO with Swin B transformer for initial asset detection.
- Combined Grounding DINO (Swin B) with YOLOv8 for enhanced MEP asset detection.
Main Results:
- Grounding DINO (Swin B) achieved high performance in open-vocabulary MEP asset detection (e.g., mIoU of 0.6586 for valves).
- The Grounding DINO (Swin B) and YOLOv8 combination demonstrated superior results, reaching mAP50 of 0.928 for valves and 0.778 for pumps.
- Performance was validated against fine-tuned models and fully supervised detectors.
Conclusions:
- The proposed method effectively addresses the limitations of existing object detection approaches for complex industrial assets.
- Leveraging vision language models and ensemble techniques offers a promising direction for open-vocabulary detection in challenging environments.
Related Concept Videos
Vision
Levels of Use of a GIS