Related Experiment Video
Updated: Jun 17, 2026

07:53
Three-Dimensional Reconstruction for the Whole Lung with Early Multiple Pulmonary Nodules
Published on: October 13, 2023
Robust artificial intelligence frameworks for lung cancer subtyping and malignancy detection on thoracic CT
Mustafa Tahsin Yilmaz1,2, Enes Algul3, Ishak Pacal4,5,6
1Department of Industrial Engineering, Faculty of Engineering, King Abdulaziz University, Jeddah, Saudi Arabia.
Scientific Reports
|June 15, 2026
Summary
Vision Transformers and Convolutional Neural Networks show promise in classifying thoracic malignancies from CT scans. Attention-based models excel at subtype discrimination, while CNNs and transformers perform well on localized abnormalities.
Area of Science:
- Radiology and Medical Imaging
- Artificial Intelligence in Oncology
- Computer Vision
Background:
- Accurate characterization of thoracic malignancies on computed tomography (CT) is challenging due to diverse visual cues needed for histological subtyping and malignancy assessment.
- Distinguishing between histological subtypes and assessing nodule malignancy requires analysis of both localized structural abnormalities and broader contextual patterns.
Purpose of the Study:
- To conduct a controlled comparative evaluation of Convolutional Neural Network (CNN) and Vision Transformer (ViT) architectures for thoracic oncology tasks.
- To assess model performance on four-class histological subtyping and three-class lung cancer classification using publicly available CT datasets.
Main Methods:
- Evaluated 30 contemporary CNN and ViT architectures on the Chest CT-Scan Images and IQ-OTH/NCCD datasets.
- Assessed performance using accuracy, precision, recall, F1-score, AUC, parameter count, confusion matrices, and Grad-CAM visualizations.
- Utilized unified experimental settings for a direct comparison of model capabilities.
Main Results:
- For histological subtyping, DeiT3-Base (ViT) achieved the highest F1-score (0.9732), outperforming CNN models like InceptionNeXt-Base.
- For lung cancer classification, EfficientNetV2-Small (CNN), Swin-Base (ViT), and MViTv2-Base (ViT) achieved the highest F1-score (0.9816).
- Grad-CAM visualizations confirmed that high-performing models focused on relevant lesion and parenchymal regions.
Conclusions:
- Attention-based architectures (ViTs) show advantages in subtype discrimination requiring contextual modeling.
- Efficient CNNs and hierarchical transformers remain competitive for tasks involving localized structural abnormalities.
- Both CNNs and ViTs demonstrate significant potential for improving thoracic malignancy characterization on CT scans.