Related Experiment Video
Updated: Jan 25, 2026

A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images
Published on: April 21, 2023
St-Swin TransNet: a spatiotemporal swin transformer-based network for self-supervised depth estimation in
Derong Yu1, Wenyuan Sun1, Junchen Wang2
1Institute of Medical Robotics, Shanghai Jiao Tong University, Dongchuan Road, Shanghai, 200240, China.
This study introduces a novel spatiotemporal Swin transformer network (ST-Swin TransNet) for accurate depth estimation in stereoscopic laparoscopic videos, improving surgical navigation. The method significantly outperforms existing approaches in depth estimation and augmented reality navigation tasks.
Area of Science:
- Computer-assisted intervention
- Medical imaging analysis
- Surgical robotics
Background:
- Depth estimation from stereoscopic laparoscopic videos is crucial for computer-assisted interventions and surgical navigation.
- Existing methods often overlook temporal information present in stereoscopic laparoscopic videos, limiting accuracy.
- Accurate depth perception is vital for enhancing surgical precision and safety.
Purpose of the Study:
- To develop a novel deep learning network for accurate depth estimation from stereoscopic laparoscopic videos.
- To leverage spatiotemporal features within video sequences for improved depth prediction.
- To evaluate the network's performance in depth estimation and its application in augmented reality surgical navigation.
Main Methods:
- A spatiotemporal Swin (ST-Swin) transformer-based network (ST-Swin TransNet) utilizing a symmetric encoder-decoder architecture was proposed.
- The network employs 12 ST-Swin blocks designed to capture spatiotemporal features via self-attention mechanisms.
- Hierarchical spatiotemporal features from binocular laparoscopic videos are exploited to predict disparity maps for depth estimation.
Main Results:
- The ST-Swin TransNet achieved a mean absolute depth error (mADE) of 3.33 mm in depth estimation tasks.
- The method demonstrated a mean absolute distance (mAD) of 1.07 mm in video see-through augmented reality (VST-AR) navigation.
- Experiments on two public datasets confirmed the superior performance of the proposed method over state-of-the-art techniques.
Conclusions:
- A novel spatiotemporal Swin transformer network (ST-Swin TransNet) was successfully developed for self-supervised depth estimation in binocular laparoscopic videos.
- The proposed method significantly outperforms existing state-of-the-art approaches in depth estimation accuracy.
- ST-Swin TransNet shows promise for enhancing computer-assisted interventions and surgical navigation systems.
Related Concept Videos
Bacterial Transformation
Griffith made an unexpected discovery when he killed the pathogenic strain and mixed its remains with the live, non-pathogenic strain. Not only did the mixture kill host mice, but it also contained living pathogenic bacteria that...
Protein Networks
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
What are Estimates?
The estimate for the mean of a sample is denoted by ͞x, whereas the mean of the population is designated as μ. Further, parameters such...
Network Covalent Solids
To break or to melt a covalent network solid, covalent bonds must be broken. Because covalent bonds are relatively strong, covalent network solids are typically...
Uniform Depth Channel Flow
Depth Perception and Spatial Vision

