Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Masking and Demasking Agents01:19

Masking and Demasking Agents

2.4K
EDTA titrations may necessitate masking and demasking agents to temporarily protect a particular metal ion in a mixture from the EDTA reaction. These agents facilitate the sequential analysis of the metal ions by forming stable complexes with some—but not all—metal ions during certain steps.
There are many masking agents, such as cyanide, fluoride, triethanolamine, thiourea, and 2,3-bis(sulfanyl)propan-1-ol (formerly 2,3-dimercapto-1-propanol), with the masking agent chosen based on...
2.4K
Modeling and Similitude01:12

Modeling and Similitude

247
Scaled modeling is a fundamental technique in engineering, enabling the study of large and complex systems by creating smaller, manageable replicas that recreate critical characteristics of the original. In hydrology and civil infrastructure, for example, scaled models of dams help analyze water flow, turbulence, and pressure. This method allows for accurate predictions of real-world behavior within a controlled environment, significantly reducing the cost and time involved in full-scale...
247
Depth Perception and Spatial Vision01:15

Depth Perception and Spatial Vision

601
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
601
Improving Translational Accuracy02:07

Improving Translational Accuracy

9.3K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
9.3K
Three-Compartment Open Model01:06

Three-Compartment Open Model

156
The three-compartment open model is a pharmacokinetic model used to describe the distribution and elimination of drugs following extravascular administration. It comprises a central compartment representing the plasma and two peripheral compartments. The highly perfused peripheral compartment represents organs and tissues with a rich blood supply, such as the liver, kidneys, and lungs. The scarcely perfused peripheral compartment represents tissues with lower blood supply, such as adipose...
156
Relative Motion Analysis using Rotating Axes-Problem Solving01:29

Relative Motion Analysis using Rotating Axes-Problem Solving

390
Consider a crane whose telescopic boom rotates with an angular velocity of 0.04 rad/s and angular acceleration of 0.02 rad/s2. Along with the rotation, the boom also extends linearly with a uniform speed of 5 m/s. The extension of the boom is measured at point D, which is measured with respect to the fixed point C on the other end of the boom. For the given instant, the distance between points C and D is 60 meters.
Here, in order to determine the magnitude of velocity and acceleration for point...
390

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Machine learning-based identification of the risk factors for postoperative nausea and vomiting in adults.

PloS one·2024
Same author

JPEG Image Enhancement with Pre-Processing of Color Reduction and Smoothing.

Sensors (Basel, Switzerland)·2023
Same author

Spectrum Correction Using Modeled Panchromatic Image for Pansharpening.

Journal of imaging·2021
Same author

Graph Neural Networks with Multiple Feature Extraction Paths for Chemical Property Estimation.

Molecules (Basel, Switzerland)·2021
Same author

Text Detection Using Multi-Stage Region Proposal Network Sensitive to Text Scale.

Sensors (Basel, Switzerland)·2021
Same author

Molecular cloning and heterologous expression of an acid-stable endoxylanase gene from Penicillium oxalicum in Trichoderma reesei.

Journal of microbiology and biotechnology·2013

Related Experiment Video

Updated: Jun 10, 2025

Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography
04:48

Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography

Published on: November 30, 2022

2.7K

TAMC: Textual Alignment and Masked Consistency for Open-Vocabulary 3D Scene Understanding.

Juan Wang1, Zhijie Wang2, Tomo Miyazaki1

  • 1Department of Communications Engineering, Graduate School of Engineering, Tohoku University, Sendai 9808579, Japan.

Sensors (Basel, Switzerland)
|October 16, 2024
PubMed
Summary

This study introduces a novel approach for 3D scene understanding using point cloud data. By employing masked consistency training and pseudo-text supervision, the method enhances environmental perception for applications like virtual reality and robotics.

Keywords:
3D Scene UnderstandingMasked ConsistencyTextual Alignmentcontrastive learningmulti-modal learningopen vocabulary

More Related Videos

Author Spotlight: Insights into the Analysis of Human Interaction with 3D Virtual Objects
06:36

Author Spotlight: Insights into the Analysis of Human Interaction with 3D Virtual Objects

Published on: October 18, 2024

897
Robotized Testing of Camera Positions to Determine Ideal Configuration for Stereo 3D Visualization of Open-Heart Surgery
05:12

Robotized Testing of Camera Positions to Determine Ideal Configuration for Stereo 3D Visualization of Open-Heart Surgery

Published on: August 12, 2021

2.0K

Related Experiment Videos

Last Updated: Jun 10, 2025

Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography
04:48

Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography

Published on: November 30, 2022

2.7K
Author Spotlight: Insights into the Analysis of Human Interaction with 3D Virtual Objects
06:36

Author Spotlight: Insights into the Analysis of Human Interaction with 3D Virtual Objects

Published on: October 18, 2024

897
Robotized Testing of Camera Positions to Determine Ideal Configuration for Stereo 3D Visualization of Open-Heart Surgery
05:12

Robotized Testing of Camera Positions to Determine Ideal Configuration for Stereo 3D Visualization of Open-Heart Surgery

Published on: August 12, 2021

2.0K

Area of Science:

  • Computer Vision
  • Robotics
  • Virtual Reality

Background:

  • Three-dimensional (3D) scene understanding relies on analyzing point cloud data for environmental perception.
  • Existing methods align 2D image features with 3D point cloud features for open-vocabulary understanding.
  • Current approaches face challenges with sparse/incomplete 3D data and lack direct text supervision during training.

Purpose of the Study:

  • To address deficiencies in 3D feature extraction from sparse/incomplete point clouds.
  • To improve the consistency between training and inference stages through direct text supervision.
  • To enhance open-vocabulary 3D scene understanding capabilities.

Main Methods:

  • Implemented a Masked Consistency training policy to handle sparse 3D features by masking parts of the data.
  • Generated pseudo-text labels by creating scene descriptions from fused 2D image descriptions.
  • Simultaneously aligned 2D-3D features and 3D-text features during the training process.

Main Results:

  • The proposed method effectively addresses the challenges of sparse and incomplete 3D point cloud data.
  • Direct text supervision during training improved the model's understanding and consistency.
  • Experimental results demonstrate superior performance compared to state-of-the-art approaches.

Conclusions:

  • The novel approach significantly advances 3D scene understanding by overcoming limitations of existing methods.
  • The combination of masked consistency and pseudo-text supervision offers a robust solution for real-world scenarios.
  • This work provides a more effective method for environmental perception in virtual reality, robotics, and beyond.