St-Swin TransNet:一个基于时空变压器的网络,用于在立体外科视频中进行自我监督的深度估计
Derong Yu1, Wenyuan Sun1, Junchen Wang2
1Institute of Medical Robotics, Shanghai Jiao Tong University, Dongchuan Road, Shanghai, 200240, China.
International journal of computer assisted radiology and surgery
|January 24, 2026
概括
这项研究引入了一种新的时空旋转变压器网络 (ST-Swin TransNet),用于在立体腹腔镜视频中准确的深度估计,改善外科导航. 该方法在深度估计和增强现实导航任务方面明显优于现有的方法.
科学领域:
- 计算机辅助干预是指计算机辅助干预.
- 医学成像分析分析 医学成像分析
- 手术机器人手术机器人手术机器人手术机器人
背景情况:
- 从立体腹腔镜视频进行深度估计对于计算机辅助干预和外科导航至关重要.
- 现有的方法往往忽略了立体腹腔镜视频中存在的时间信息,限制了准确性.
- 准确的深度感知对于提高手术精度和安全至关重要.
研究的目的:
- 开发一个新的深度学习网络,从立体腹腔镜视频中准确地估计深度.
- 为了利用视频序列中的时空特征来改进深度预测.
- 评估网络的深度估计性能及其在增强现实外科导航中的应用.
主要方法:
- 提出了一个基于变压器的空间时空Swin (ST-Swin) 网络 (ST-Swin TransNet) 使用对称的编码器解码器架构.
- 该网络使用12个ST-Swin块,旨在通过自我注意机制捕捉时空特征.
- 来自双眼腹腔镜视频的层次空间时空特征被利用来预测差异地图用于深度估计.
主要成果:
- 在深度估计任务中,ST-Swin TransNet 实现了平均绝对深度误差 (mADE) 3.33 mm.
- 该方法在视频透视增强现实 (VST-AR) 导航中显示了1.07mm的平均绝对距离 (mAD).
- 在两个公共数据集上的实验证实了拟议方法在最先进的技术上具有优越的性能.
结论:
- 一个新的时空Swin变压器网络 (ST-Swin TransNet) 已成功开发,用于双眼腹腔镜视频中自我监督的深度估计.
- 拟议的方法在深度估计准确性方面明显优于现有的最先进的方法.
- ST-Swin TransNet显示了增强计算机辅助干预和外科导航系统的前景.
相关概念视频
Bacterial Transformation
59.5K
In 1928, bacteriologist Frederick Griffith worked on a vaccine for pneumonia, which is caused by Streptococcus pneumoniae bacteria. Griffith studied two pneumonia strains in mice: one pathogenic and one non-pathogenic. Only the pathogenic strain killed host mice.
Griffith made an unexpected discovery when he killed the pathogenic strain and mixed its remains with the live, non-pathogenic strain. Not only did the mixture kill host mice, but it also contained living pathogenic bacteria that...
Griffith made an unexpected discovery when he killed the pathogenic strain and mixed its remains with the live, non-pathogenic strain. Not only did the mixture kill host mice, but it also contained living pathogenic bacteria that...
59.5K
Protein Networks
4.5K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
4.5K
What are Estimates?
8.2K
It isn't easy to measure a parameter such as the mean height or the mean weight of a population. So, we draw samples from the population and calculate the mean height or mean weight of the individuals in the sample. This sample data acts as a representative measure of the population parameter. These sample statistics are known as estimates.
The estimate for the mean of a sample is denoted by ͞x, whereas the mean of the population is designated as μ. Further, parameters such...
The estimate for the mean of a sample is denoted by ͞x, whereas the mean of the population is designated as μ. Further, parameters such...
8.2K
Network Covalent Solids
16.1K
Network covalent solids contain a three-dimensional network of covalently bonded atoms as found in the crystal structures of nonmetals like diamond, graphite, silicon, and some covalent compounds, such as silicon dioxide (sand) and silicon carbide (carborundum, the abrasive on sandpaper). Many minerals have networks of covalent bonds.
To break or to melt a covalent network solid, covalent bonds must be broken. Because covalent bonds are relatively strong, covalent network solids are typically...
To break or to melt a covalent network solid, covalent bonds must be broken. Because covalent bonds are relatively strong, covalent network solids are typically...
16.1K
Uniform Depth Channel Flow
540
Uniform depth channel flow keeps fluid depth consistent along channels such as irrigation canals. In natural channels, such as rivers, approximate uniform flow is often assumed. This condition occurs when the channel’s bottom slope matches the energy slope, balancing potential energy lost from gravity with head loss due to shear stress. This balance prevents depth changes along the channel length, resulting in a steady, uniform flow.Uniform flow in open channels with a constant cross-section...
540
Depth Perception and Spatial Vision
1.9K
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
1.9K


