Related Experiment Video
Updated: Jul 16, 2026

05:12
Robotized Testing of Camera Positions to Determine Ideal Configuration for Stereo 3D Visualization of Open-Heart Surgery
Published on: August 12, 2021
Confidence-Guided Fusion for Self-Supervised Monocular Depth Estimation in Endoscopy
Shuang Li1, Hongbo Wang2, Zhaoxu Hu2
1School of Information and Artificial Intelligence, Xi'an University of Architecture and Technology Huaqing College, Xi'an 710043, China.
Sensors (Basel, Switzerland)
|July 15, 2026
Summary
CoDepth fuses diffusion and discriminative models for enhanced monocular depth estimation (MDE) in surgery. This confidence-guided approach improves accuracy and robustness, aiding surgical navigation.
Area of Science:
- Computer Vision
- Medical Imaging
- Surgical Technology
Background:
- Accurate monocular depth estimation (MDE) is crucial for endoscopic surgery, improving navigation and depth perception.
- Existing diffusion and discriminative models have complementary strengths but also asymmetric error patterns.
- Discriminative models excel at geometric boundaries but fail in homogeneous areas; diffusion models capture texture but may lack structural coherence.
Purpose of the Study:
- To develop a novel framework, CoDepth, that harmonizes heterogeneous depth estimators by leveraging confidence-guided fusion.
- To systematically exploit the complementary strengths of diffusion-based and discriminative depth estimation models.
- To improve the accuracy, robustness, and generalizability of MDE in endoscopic surgery.
Main Methods:
- Introduced CoDepth, a framework employing confidence-guided fusion to integrate outputs from diverse depth estimators.
- Developed a complementary map extractor to identify disparity disagreements between models.
- Integrated a cross-attention module for context-aware feature integration and a probabilistic confidence network for adaptive fusion weights.
Main Results:
- CoDepth demonstrated improved overall performance on the SCARED dataset compared to single-model baselines, particularly in Abs Rel and δ-based accuracy.
- The framework showed encouraging cross-domain generalization, achieving competitive performance on SERV-CT, Hamlyn, and C3VD datasets without fine-tuning.
- CoDepth exhibited enhanced robustness against synthetic corruptions like low-light, Gaussian noise, and impulse noise.
Conclusions:
- Confidence-guided complementary fusion offers a practical integration paradigm for combining heterogeneous endoscopic depth estimators.
- CoDepth enhances MDE accuracy and robustness, showing significant potential for improving surgical navigation and clinical utility.
- The framework's cross-domain generalization and robustness highlight its adaptability to complex and varied surgical environments.
Related Concept Videos
Depth Perception and Spatial Vision
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
Endoscopic Procedures III: Video Capsule Endoscopy
Capsule endoscopy, or wireless or video capsule endoscopy, is a diagnostic procedure for examining the entire gastrointestinal tract. Patients swallow a capsule about the size of a vitamin tablet. The capsule is equipped with a transmitter, a battery, an LED light source, and a color video camera to capture images throughout the gastrointestinal tract. This procedure is particularly useful for diagnosing conditions such as Crohn's disease, ulcerative colitis, tumors, polyps, ulcers, unexplained...