Related Experiment Video
Updated: Mar 10, 2026

Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography
Published on: November 30, 2022
Two-stage deep learning framework for laterally spreading tumors detection using self-supervised learning and
Menghui Wang1, Zhanpeng Shi2, Yiwen Wang3
1Department of Gastroenterology, Shanghai General Hospital, Shanghai Jiao Tong University School of Medicine, Shanghai, China.
Background:
Laterally spreading tumors (LSTs) of the colorectum are flat lesions with precancerous potential that are often overlooked during routine colonoscopy due to their subtle morphology and low prevalence, posing a significant challenge for early detection.
Aim:
To address the scarcity of labeled LSTs data and the limitations of conventional supervised learning, we formulate LSTs recognition as an image-level binary classification problem (LSTs vs non-LSTs) using colonoscopy images. The aim of this study is to achieve reliable discrimination under limited expert annotations.
Methods:
A large-scale dataset of 150,168 colonoscopy images was retrospectively collected from 12,376 patients at Shanghai General Hospital between July 2021 and July 2025. The framework included two components: (1) DINO self-supervised pretraining on 150,168 unlabeled images to learn robust visual representations, and (2) Prototypical Networks performing few-shot classification with 2799 labeled training images. The model was evaluated on 601 test images using comprehensive metrics including area under the receiver operating characteristic curve (ROC-AUC), sensitivity, and specificity.
Results:
The proposed model demonstrated robust performance, achieving an overall accuracy of 72.4% (95% confidence interval (CI): 67.2-77.7%), sensitivity of 74.7%, specificity of 70.2%, and ROC-AUC of 0.798 (95% CI: 0.746-0.850). Notably, the model also attained an F1-score of 0.730 and a precision-recall AUC of 0.828. The optimal decision threshold determined by Youden's index was 0.338. With a rapid processing speed of 50 ms per image, the framework is well-suited for real-time clinical applications.
Conclusions:
By combining self-supervised representation learning with few-shot classification, our framework reduces the need for large annotated datasets while maintaining clinically relevant performance. This approach offers a practical pathway toward AI-assisted LSTs detection in resource-constrained endoscopic settings and lays the groundwork for future integration into real-time colonoscopy support systems.