Related Experiment Video
Updated: Jan 13, 2026

10:26
Author Spotlight: A 3D Digital Model for the Diagnosis and Treatment of Pulmonary Nodules
Published on: May 19, 2023
2.5K
Vision-language model-based semantic-guided imaging biomarker for lung nodule malignancy prediction.
Luoting Zhuang1, Seyed Mohammad Hossein Tabatabaei1, Ramin Salehi-Rad2
1Medical & Imaging Informatics, Department of Radiological Sciences, David Geffen School of Medicine at UCLA, Los Angeles, 90095, CA, USA.
Journal of Biomedical Informatics
|October 29, 2025
Summary
This study introduces a novel machine learning approach for lung cancer prediction using radiologist-derived semantic features and a vision-language model. The method enhances diagnostic accuracy and interpretability in lung nodule analysis.
Area of Science:
- Artificial Intelligence in Medical Imaging
- Radiology and Oncology
- Computer-Aided Diagnosis
Background:
- Machine learning models for lung nodule malignancy assessment often struggle with interpretability and reliance on manual annotation.
- Existing models face challenges in robustness due to sensitivity to imaging variations in clinical settings.
Purpose of the Study:
- To develop a machine learning model that integrates semantic features from radiologists to predict lung cancer.
- To enhance the clinical relevance, robustness, and explainability of lung nodule malignancy assessment.
Main Methods:
- Utilized 938 low-dose CT scans (NLST) and 2,625 lesions (LIDC) with semantic features; external datasets from UCLA, LUNGx, and Duke were included.
- Converted structured semantic features into text and fine-tuned a pretrained Contrastive Language-Image Pre-training (CLIP) model using parameter-efficient fine-tuning.
- Predicted one-year lung cancer diagnosis by aligning imaging and semantic text features.
Main Results:
- Achieved an AUROC of 0.901 and AUPRC of 0.776 on the NLST test set, outperforming SOTA models.
- Demonstrated robust performance across multiple external validation datasets.
- Zero-shot inference using CLIP yielded high AUROCs for nodule margin (0.807), consistency (0.812), and pleural attachment (0.840).
Conclusions:
- The proposed vision-language model integrating semantic features surpasses SOTA models in lung cancer prediction from diverse CT scans.
- The approach offers explainable outputs, assisting clinicians in understanding model predictions.
- The study provides a pathway for more robust and interpretable AI in lung cancer screening.

