在内镜视频中利用时空线索进行自我监督的立体深度估计
IEEE transactions on medical imaging
|January 29, 2026
概括
这项研究引入了一种新的语自我监督框架,用于内镜深度估计,通过利用时间和立体信息,显著提高动态手术视频的准确性和减少人工物.
科学领域:
- 医疗成像医学成像
- 计算机视觉 计算机视觉
- 手术技术 手术技术
背景情况:
- 当前自主监督的内镜深度估计方法往往忽视时间动态,导致文物.
- 这种限制在动态的外科环境中阻碍了准确的深度重建.
研究的目的:
- 开发一个强大的自我监督的框架,用于内镜视频中准确的深度估计.
- 通过结合时间和立体信息来解决现有方法的局限性.
主要方法:
- 引入了一个采用立体声和时间数据的语自我监督框架.
- 开发了可变形时空交叉视图融合 (DSCF) 和多尺度选择性时空聚合 (MSSA) 模块.
- 实施了为有效的边缘部署和单眼深度估计而建立的三阶段培训协议.
主要成果:
- 在四个公共内镜数据集 (SCARED,SERV-CT,EndoNeRF,Hamlyn) 上实现了最先进的性能.
- 在深度准确度和图像重建质量方面表现出显著的改进.
- 与现有方法相比,根平均平方误差 (RMSE) 减少了超过9.1%.
结论:
- 拟议的框架有效地通过整合时空信息来增强内镜深度估计.
- 该方法为实时手术应用和边缘设备部署提供了一个有前途的解决方案.
相关概念视频
Endoscopic Procedures III: Video Capsule Endoscopy
835
Capsule endoscopy, or wireless or video capsule endoscopy, is a diagnostic procedure for examining the entire gastrointestinal tract. Patients swallow a capsule about the size of a vitamin tablet. The capsule is equipped with a transmitter, a battery, an LED light source, and a color video camera to capture images throughout the gastrointestinal tract. This procedure is particularly useful for diagnosing conditions such as Crohn's disease, ulcerative colitis, tumors, polyps, ulcers,...
835
Non-Verbal Cues
329
Non-verbal communication extends beyond gestures and facial expressions to include vocal elements known as paralanguage. Paralanguage consists of non-verbal vocal cues such as pitch, loudness, speech rate, pauses, and non-verbal vocalizations like laughter, sighs, and moans. These elements not only accompany speech but also provide critical emotional and contextual information.The Role of Paralanguage in CommunicationParalanguage adds depth to spoken language by conveying emotions and...
329
What are Estimates?
8.8K
It isn't easy to measure a parameter such as the mean height or the mean weight of a population. So, we draw samples from the population and calculate the mean height or mean weight of the individuals in the sample. This sample data acts as a representative measure of the population parameter. These sample statistics are known as estimates.
The estimate for the mean of a sample is denoted by ͞x, whereas the mean of the population is designated as μ. Further, parameters such...
The estimate for the mean of a sample is denoted by ͞x, whereas the mean of the population is designated as μ. Further, parameters such...
8.8K
Uniform Depth Channel Flow
582
Uniform depth channel flow keeps fluid depth consistent along channels such as irrigation canals. In natural channels, such as rivers, approximate uniform flow is often assumed. This condition occurs when the channel’s bottom slope matches the energy slope, balancing potential energy lost from gravity with head loss due to shear stress. This balance prevents depth changes along the channel length, resulting in a steady, uniform flow.Uniform flow in open channels with a constant cross-section...
582
Depth Perception and Spatial Vision
2.0K
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
2.0K
Estimation of k and VD of Aminoglycosides
240
Aminoglycosides are a class of antibiotics used to treat various bacterial infections. Clinicians must determine the elimination rate constant (k) and volume of distribution (VD) to optimize therapeutic efficacy and minimize toxicity. The k value represents the rate at which the drug is removed from the body, and the VD reflects the degree to which the drug distributes into body tissues. Accurately estimating these parameters allows healthcare professionals to tailor drug dosing to individual...
240


