Related Experiment Video
Updated: May 6, 2026

A Swin Transformer-Based Model for Thyroid Nodule Detection in Ultrasound Images
Published on: April 21, 2023
Trackerless 3D ultrasound volume reconstruction from 2D freehand scans using a hybrid transformer-CNN framework
Wenfeng He1, Richard L J Qiu1, Chulong Zhang2
1Department of Radiation Oncology and Winship Cancer Institute, Emory University, Atlanta, Georgia, USA.
Background:
Generating three-dimensional (3D) ultrasound (US) data from conventional two-dimensional (2D) freehand acquisitions typically necessitates external tracking systems to ascertain probe positioning. However, these hardware-based solutions often present practical challenges, including substantial cost, increased setup complexity, and susceptibility to interference or line-of-sight issues, thereby limiting their widespread clinical integration. A tracker-free approach to 3D US reconstruction could dramatically improve the accessibility and applicability of volumetric ultrasound in standard medical procedures.
Purpose:
This study presents a method for 3D US reconstruction from 2D B-mode image sequences that eliminates the need for external tracking. The goal is to improve reconstruction accuracy and usability for diagnostic, preoperative, and intraoperative applications.
Methods:
A hybrid Transformer-convolutional neural network (CNN) is trained end-to-end to regress inter-frame six degrees of freedom (6-DoF) poses. Global self-attention captures long-range probe motion, whereas convolutional layers refine local speckle patterns. Experiments used 53 forearm scans (Dataset 1) and 36 prostate scans (Dataset 2). Reconstruction performance was assessed with Dice Similarity Coefficient (DSC) and trajectory drift (Dataset 1) plus structural similarity (SSIM) and peak signal-to-noise ratio (PSNR) (Dataset 2).
Results:
In Dataset 1, the proposed model achieved a DSC of 0.71 0.19 and a drift of 13.25 9.17 , outperforming the 2D CNN (0.62 0.26, 18.92 11.20 ) and ConvLSTM (0.67 0.32, 18.36 5.44 ). In Dataset 2, it obtained an SSIM of 0.845 and a PSNR of 29.10 dB, exceeding the 2-D CNN (0.742, 26.40 dB) and ConvLSTM (0.812, 27.85 dB). All improvements were statistically significant as determined by paired t-tests (p 0.01).
Conclusions:
The tracker-free Transformer-CNN consistently improves volumetric overlap, trajectory stability, and image quality relative to established CNN- or recurrent neural network (RNN)-based schemes, demonstrating a practical, cost-effective route to high-fidelity 3D ultrasound across diverse clinical protocols.

