Related Experiment Video
Updated: Oct 16, 2025

Author Spotlight: A 3D Digital Model for the Diagnosis and Treatment of Pulmonary Nodules
Published on: May 19, 2023
Detection of Pneumothorax with Deep Learning Models: Learning From Radiologist Labels vs Natural Language Processing
James Thomas Patrick Decourcy Hallinan1, Mengling Feng2, Dianwen Ng3
1Department of Diagnostic Imaging, National University Hospital, Singapore.
Deep learning models for pneumothorax detection performed better when trained with radiologist labels compared to natural language processing (NLP) labels. This indicates that expert-annotated data improves model accuracy and generalizability for medical imaging tasks.
Area of Science:
- Medical Imaging
- Artificial Intelligence
- Radiology
Background:
- Pneumothorax detection in chest radiographs is crucial for patient care.
- Deep learning models offer potential for automated detection.
- Labeling datasets for training these models can be time-consuming and resource-intensive.
Purpose of the Study:
- To compare the performance of deep learning models for pneumothorax detection.
- To evaluate models trained with radiologist-annotated labels versus natural language processing (NLP)-derived labels.
- To assess model generalizability using internal and external test sets.
Main Methods:
- Utilized the NIH ChestX-ray14 dataset, comprising over 112,000 chest radiographs.
- Created two datasets: one with NLP-derived pneumothorax labels and another with radiologist-confirmed labels.
- Trained three convolutional neural network (CNN) architectures (ResNet-50, DenseNet-121, EfficientNetB3) independently on both datasets.
- Evaluated model performance using the area under the receiver operating characteristic curve (AUC) on both the NIH internal test set and an external emergency department test set.
Main Results:
- CNN models trained with radiologist labels consistently achieved significantly higher AUCs compared to those trained with NLP labels across all architectures on both internal and external test sets.
- For instance, on the external test set, AUCs for radiologist-trained models ranged from 0.806 to 0.915, outperforming NLP-trained models (AUCs 0.686 to 0.822).
- The most significant performance improvement was observed with the EfficientNetB3 architecture, demonstrating the impact of label quality on model efficacy.
Conclusions:
- Radiologist-annotated labels lead to superior performance and generalizability of deep learning models for pneumothorax detection compared to NLP-derived labels.
- The findings underscore the importance of high-quality, expert-annotated data in developing robust medical AI tools.
- These results suggest that investing in radiologist-curated datasets can enhance the reliability and clinical utility of AI-powered diagnostic systems.
More Related Videos
07:53Author Spotlight: Advancing 3D Modeling for Enhanced Diagnosis and Treatment of Pulmonary Nodules in Early-Stage Lung Cancer
Published on: October 13, 2023
08:05Lung CT Segmentation to Identify Consolidations and Ground Glass Areas for Quantitative Assesment of SARS-CoV Pneumonia
Published on: December 19, 2020
Related Concept Videos
Pneumothorax-I
Pneumothorax can be even further classified as spontaneous, traumatic, and tension pneumothorax.
Pneumothorax-II
Clinical Manifestations: