Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Vision01:24

Vision

48.6K
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
48.6K
Levels of Use of a GIS01:29

Levels of Use of a GIS

511
Geographic Information Systems (GIS) operate across three levels of application, each representing an increasing degree of complexity: data management, analysis, and prediction. These levels reflect the expanding functionality and versatility of GIS technology in handling spatial data for diverse purposes.Data ManagementAt its foundational level, GIS serves as a tool for data management, enabling the input, storage, retrieval, and organization of spatial data. This level is often employed in...
511

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Joint Function and Movement Variability During Daily Living Activities Performed Throughout the Home Setting: A Digital Twin Modeling Study.

Sensors (Basel, Switzerland)·2025
Same author

Enhancing Healthcare through Sensor-Enabled Digital Twins in Smart Environments: A Comprehensive Analysis.

Sensors (Basel, Switzerland)·2024
Same author

Region-income-based prioritisation of Sustainable Development Goals by Gradient Boosting Machine.

Sustainability science·2022
Same author

A Blockchain-Based Spatial Crowdsourcing System for Spatial Information Collection Using a Reward Distribution.

Sensors (Basel, Switzerland)·2021
Same author

Influence of pedestrian age and gender on spatial and temporal distribution of pedestrian crashes.

Traffic injury prevention·2017

Related Experiment Video

Updated: May 5, 2026

Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
08:25

Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment

Published on: May 7, 2019

8.6K

A Vision Language-Based Framework for Detecting Industrial Mechanical, Electrical, and Plumbing Assets Using

Masoud Kamali1, Behnam Atazadeh1, Abbas Rajabifard1

  • 1The Centre for Spatial Data Infrastructures and Land Administration, Department of Infrastructure Engineering, The University of Melbourne, Melbourne, VIC 3010, Australia.

Sensors (Basel, Switzerland)
|May 4, 2026
PubMed
Summary

This study introduces a novel method for detecting unseen mechanical, electrical, and plumbing (MEP) assets using vision language models and object detectors. The approach significantly improves open-vocabulary detection of complex industrial assets.

Keywords:
MEP assetsclose-set object detectorsopen-vocabulary object detectionvision language models

Related Experiment Videos

Last Updated: May 5, 2026

Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
08:25

Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment

Published on: May 7, 2019

8.6K

Area of Science:

  • Computer Vision
  • Artificial Intelligence
  • Industrial Asset Management

Background:

  • Object detection models struggle with limited training data diversity and generalizing to new asset categories in industrial settings.
  • Mechanical, electrical, and plumbing (MEP) assets present unique spatial and geometric complexities challenging current detection methods.

Purpose of the Study:

  • To develop an effective approach for detecting unseen MEP assets in industrial environments.
  • To leverage pre-trained vision language models and close-set object detectors for open-vocabulary asset detection using unlabeled data.

Main Methods:

  • Utilized pre-trained vision language models and close-set object detectors.
  • Employed Grounding DINO with Swin B transformer for initial asset detection.
  • Combined Grounding DINO (Swin B) with YOLOv8 for enhanced MEP asset detection.

Main Results:

  • Grounding DINO (Swin B) achieved high performance in open-vocabulary MEP asset detection (e.g., mIoU of 0.6586 for valves).
  • The Grounding DINO (Swin B) and YOLOv8 combination demonstrated superior results, reaching mAP50 of 0.928 for valves and 0.778 for pumps.
  • Performance was validated against fine-tuned models and fully supervised detectors.

Conclusions:

  • The proposed method effectively addresses the limitations of existing object detection approaches for complex industrial assets.
  • Leveraging vision language models and ensemble techniques offers a promising direction for open-vocabulary detection in challenging environments.