Related Experiment Video
Updated: Jun 3, 2025

05:12
Robotized Testing of Camera Positions to Determine Ideal Configuration for Stereo 3D Visualization of Open-Heart Surgery
Published on: August 12, 2021
2.0K
Residual Vision Transformer and Adaptive Fusion Autoencoders for Monocular Depth Estimation
Wei-Jong Yang1, Chih-Chen Wu2, Jar-Ferr Yang2
1Department of Artificial Intelligence and Computer Engineering, National Chin-Yi University of Technology, Taichung 411, Taiwan.
Sensors (Basel, Switzerland)
|January 11, 2025
Summary
This study introduces a novel autoencoder for monocular depth estimation, enhancing 3D sensing with a single camera. The model achieves superior accuracy and reduced error in depth map prediction, improving applications like autonomous driving.
Area of Science:
- Computer Vision
- Deep Learning
- 3D Sensing
Background:
- Monocular depth estimation is crucial for applications like 3D scene reconstruction and autonomous driving.
- Deep learning advancements have enabled monocular depth estimation to surpass traditional stereo camera systems.
- Accurate depth perception from single-view images remains a significant challenge.
Purpose of the Study:
- To propose an end-to-end supervised monocular depth estimation autoencoder using a single camera.
- To enhance the precision of depth maps through a novel network architecture.
- To improve depth estimation performance for foreground objects.
Main Methods:
- Developed an autoencoder with a mixed convolution neural network and vision transformers encoder.
- Implemented an adaptive fusion decoder for effective feature merging.
- Utilized a human perception-aligned loss function for training.
Main Results:
- The proposed autoencoder effectively predicts depth maps from single-view color images.
- Achieved a 28% increase in the first accuracy rate compared to existing methods.
- Reduced the root mean square error by approximately 27% on the NYU dataset.
Conclusions:
- The developed autoencoder demonstrates high-precision monocular depth estimation capabilities.
- The hybrid encoder and adaptive decoder architecture effectively capture multi-scale features.
- The perception-aligned loss function improves focus on critical foreground object depth.
Related Concept Videos
Depth Perception and Spatial Vision
548
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
548
Vision
52.9K
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
52.9K

