Related Experiment Video
Updated: Jun 18, 2025

04:48
Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography
Published on: November 30, 2022
2.7K
Leveraging Pretrained Transformers for Efficient Segmentation and Lesion Detection in Cone-Beam Computed Tomography
Rui Qi Chen1, Yeonju Lee1, Hao Yan2
1H. Milton Stewart School of Industrial and Systems Engineering, Georgia Institute of Technology, Atlanta, Georgia.
Journal of Endodontics
|August 3, 2024
Summary
Pretrained transformer models like Swin-UNETR show excellent performance in segmenting jaw lesions from cone-beam computed tomography (CBCT) scans. This artificial intelligence approach improves lesion detection accuracy, even with limited training data.
Area of Science:
- Dentistry
- Medical Imaging
- Artificial Intelligence
Background:
- Cone-beam computed tomography (CBCT) is crucial for detecting jaw lesions but interpretation is challenging and time-consuming.
- Artificial intelligence (AI) in CBCT segmentation can enhance lesion detection accuracy.
- Consistent automated lesion detection is difficult, particularly with limited training data.
Purpose of the Study:
- To evaluate the effectiveness of pretrained transformer-based architectures for semantic segmentation of CBCT volumes.
- To assess the application of these models in periapical lesion detection.
Main Methods:
- CBCT volumes (n=138) were annotated with labels including 'lesion,' 'restorative material,' 'bone,' 'tooth structure,' and 'background.'
- U-Net and Swin-UNETR models (pretrained and from scratch) were trained using subsets of the annotated CBCTs.
- Performance was evaluated using Sørensen-Dice coefficient (DICE) for segmentation and sensitivity/specificity for lesion detection, with varying training sample sizes (20, 40, 60, 103).
Main Results:
- The pretrained Swin-UNETR model achieved high DICE scores across all labels, including 0.8512 for 'lesion' with 103 samples.
- Lesion detection performance remained statistically similar between models trained with 103 and 60 images, with 1.00 sensitivity and 0.94 specificity using 60 images.
- Pretrained Swin-UNETR outperformed Swin-UNETR-SCRATCH and U-Net in segmentation and lesion detection specificity, especially with limited training data.
Conclusions:
- Transformer-based Swin-UNETR architectures demonstrate excellent semantic segmentation and periapical lesion detection capabilities.
- Using pretrained models offers a viable alternative to traditional U-Net architectures, requiring smaller training datasets for effective CBCT analysis.

