Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Depth Perception and Spatial Vision01:15

Depth Perception and Spatial Vision

840
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
840
Transformers with Off-Nominal Turns Ratios01:25

Transformers with Off-Nominal Turns Ratios

193
In scenarios involving parallel transformers with disparate ratings, developing per-unit models requires accommodating off-nominal turns ratios. This situation arises when the selected base voltages are not proportional to the transformer’s voltage ratings. Consider a transformer where the rated voltages are related by the term a. If the chosen voltage bases satisfy a relationship involving term b, term c is defined as the ratio of these bases. This ratio is then substituted into the...
193
Deconvolution01:20

Deconvolution

233
Deconvolution, also known as inverse filtering, is the process of extracting the impulse response from known input and output signals. This technique is vital in scenarios where the system's characteristics are unknown, and they must be inferred from the observable signals.
Deconvolution involves several mathematical techniques to derive the impulse response. One common approach is polynomial division. In this method, the input and output sequences are treated as coefficients of...
233

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Differential Image Sensor With Decoupled Static and Dynamic Outputs.

Advanced materials (Deerfield Beach, Fla.)·2026
Same author

Retraction Note: Effect of photobiological regulation of green laser on orthodontic tooth retention in rats.

Lasers in medical science·2026
Same author

How do infrared and near-infrared lasers influence autophagy and alveolar bone remodeling in rats during orthodontic tooth retention?

BMC oral health·2025
Same author

Effect of photobiological regulation of green laser on orthodontic tooth retention in rats.

Lasers in medical science·2025
Same author

Feature Separation in Diffuse Lung Disease Image Classification by Using Evolutionary Algorithm-Based NAS.

IEEE journal of biomedical and health informatics·2024
Same author

High-Performance Binocular Disparity Prediction Algorithm for Edge Computing.

Sensors (Basel, Switzerland)·2024

Related Experiment Video

Updated: Aug 25, 2025

Assessing Binocular Central Visual Field and Binocular Eye Movements in a Dichoptic Viewing Condition
07:45

Assessing Binocular Central Visual Field and Binocular Eye Movements in a Dichoptic Viewing Condition

Published on: July 21, 2020

4.5K

Transformer Based Binocular Disparity Prediction with Occlusion Predict and Novel Full Connection Layers.

Yi Liu1,2, Xintao Xu1,3, Bajian Xiang1,2

  • 1Institute of Semiconductors, Chinese Academy of Sciences, Beijing 100083, China.

Sensors (Basel, Switzerland)
|October 14, 2022
PubMed
Summary

This study introduces a novel Transformer-based algorithm for disparity prediction, overcoming limitations of convolutional neural networks. The model achieves high accuracy and efficiency in depth estimation tasks.

Keywords:
attentionbinocular disparitytransformer

More Related Videos

A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images
04:23

A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images

Published on: April 21, 2023

1.9K
How to Build a Dichoptic Presentation System That Includes an Eye Tracker
05:48

How to Build a Dichoptic Presentation System That Includes an Eye Tracker

Published on: September 6, 2017

8.6K

Related Experiment Videos

Last Updated: Aug 25, 2025

Assessing Binocular Central Visual Field and Binocular Eye Movements in a Dichoptic Viewing Condition
07:45

Assessing Binocular Central Visual Field and Binocular Eye Movements in a Dichoptic Viewing Condition

Published on: July 21, 2020

4.5K
A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images
04:23

A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images

Published on: April 21, 2023

1.9K
How to Build a Dichoptic Presentation System That Includes an Eye Tracker
05:48

How to Build a Dichoptic Presentation System That Includes an Eye Tracker

Published on: September 6, 2017

8.6K

Area of Science:

  • Computer Vision
  • Deep Learning
  • Artificial Intelligence

Background:

  • Convolutional Neural Networks (CNNs) for depth estimation face limitations in disparity range, occlusion handling, and global context perception.
  • Traditional methods struggle with acquiring accurate disparities beyond a predefined range and ensuring matching uniqueness.

Purpose of the Study:

  • To propose a Transformer-based disparity prediction algorithm that addresses the shortcomings of CNNs in depth estimation.
  • To enhance feature extraction, disparity matching, and occlusion prediction for improved depth accuracy.

Main Methods:

  • Utilized a Swin-SPP module for feature extraction based on Swin Transformer.
  • Developed a Transformer disparity matching network employing self-attention and cross-attention mechanisms.
  • Integrated an occlusion prediction sub-network and a double skip connection fully connected layer to improve training stability and inference accuracy.

Main Results:

  • Achieved an EPE (Absolute Error) of 0.57 on KITTI 2012 and 0.61 on KITTI 2015.
  • Obtained a 3PE (Percentage Error > 3px) of 1.74% on KITTI 2012 and 1.56% on KITTI 2015.
  • Demonstrated efficient inference time (0.46s) with a low parameter count (2.6M).

Conclusions:

  • The proposed Transformer-based model significantly outperforms existing algorithms in depth estimation accuracy and efficiency.
  • The novel architecture effectively handles limitations related to disparity range, occlusion, and global context.
  • The algorithm shows great advantages in various evaluation metrics for stereo depth estimation.