Related Experiment Video
Updated: Jan 25, 2026

A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images
Published on: April 21, 2023
St-Swin TransNet: a spatiotemporal swin transformer-based network for self-supervised depth estimation in
Derong Yu1, Wenyuan Sun1, Junchen Wang2
1Institute of Medical Robotics, Shanghai Jiao Tong University, Dongchuan Road, Shanghai, 200240, China.
Purpose:
Depth estimation from stereoscopic laparoscopic videos is of vital importance in computer-assisted intervention due to its potential for downstream tasks in laparoscopic surgical navigation. Previous works mostly focus on depth estimation from static frames, while temporal information in stereoscopic laparoscopic videos is largely ignored.
Methods:
A spatiotemporal swin (ST-Swin) transformer-based network, referred to as ST-Swin TransNet, is proposed for depth estimation in stereoscopic surgical videos. Built upon a symmetric encoder-decoder architecture consisting of 12 ST-Swin blocks, ST-Swin TransNet extracts spatiotemporal features for efficient and accurate depth estimation, where the ST-Swin blocks are designed to capture spatiotemporal information from stereo video sequences via self-attention mechanism. Given binocular laparoscopic videos, ST-Swin TransNet exploits hierarchical spatiotemporal features to predict disparity maps.
Results:
Comprehensive experiments are conducted on two typical yet challenging public datasets to evaluate the performance of the proposed method. We additionally demonstrate the feasibility of applying ST-Swin TransNet to video see-through augmented reality (VST-AR) navigation in laparoscopic surgery. Our method achieved a mean absolute depth error (mADE) of 3.33 mm in depth estimation and a mean absolute distance (mAD) of 1.07 mm in VST-AR navigation.
Conclusion:
A spatiotemporal swin transformer-based network for self-supervised depth estimation in binocular laparoscopic surgical videos was developed. Results from the comprehensive experiments demonstrate the superior performance of the proposed method over the state-of-the-art methods.
Related Concept Videos
Bacterial Transformation
Griffith made an unexpected discovery when he killed the pathogenic strain and mixed its remains with the live, non-pathogenic strain. Not only did the mixture kill host mice, but it also contained living pathogenic bacteria that...
Protein Networks
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
What are Estimates?
The estimate for the mean of a sample is denoted by ͞x, whereas the mean of the population is designated as μ. Further, parameters such...
Network Covalent Solids
To break or to melt a covalent network solid, covalent bonds must be broken. Because covalent bonds are relatively strong, covalent network solids are typically...
Uniform Depth Channel Flow
Depth Perception and Spatial Vision

