Related Experiment Video
Updated: Sep 11, 2025

10:26
Author Spotlight: A 3D Digital Model for the Diagnosis and Treatment of Pulmonary Nodules
Published on: May 19, 2023
2.0K
Vision-Language Model-Based Semantic-Guided Imaging Biomarker for Lung Nodule Malignancy Prediction.
Luoting Zhuang1, Seyed Mohammad Hossein Tabatabaei1, Ramin Salehi-Rad2
1Medical & Imaging Informatics, Department of Radiological Sciences, David Geffen School of Medicine at UCLA, Los Angeles, 90095, CA, USA.
Arxiv
|August 13, 2025
Summary
This study introduces a novel machine learning approach using radiologist-guided semantic features to improve lung cancer prediction from CT scans. The model achieves superior accuracy and provides explainable results, enhancing clinical applicability.
Area of Science:
- Artificial Intelligence in Radiology
- Machine Learning for Medical Imaging
- Lung Cancer Detection and Diagnosis
Background:
- Current machine learning models for lung nodule malignancy assessment often rely on manual annotations and lack interpretability.
- These models are sensitive to imaging variations, limiting their real-world clinical utility.
- There is a need for robust, explainable AI models that integrate clinical insights for lung cancer prediction.
Purpose of the Study:
- To develop and validate a machine learning model that integrates semantic features from radiologists' assessments with imaging data.
- To enhance the clinical relevance, robustness, and interpretability of lung cancer prediction models.
- To guide the model in learning clinically meaningful imaging features for accurate lung cancer diagnosis.
Main Methods:
- Utilized low-dose CT scans from the National Lung Screening Trial (NLST) and other external datasets (LIDC, UCLA, LUNGx, Duke).
- Extracted 2D nodule slices and converted structured semantic features into sentences using Gemini.
- Fine-tuned a pre-trained Contrastive Language-Image Pre-training (CLIP) model using parameter-efficient fine-tuning to align imaging and semantic features for one-year lung cancer prediction.
Main Results:
- The proposed model achieved superior performance on the NLST test set, with an Area Under the Receiver Operating Characteristic Curve (AUROC) of 0.901 and Area Under the Precision-Recall Curve (AUPRC) of 0.776.
- Demonstrated robust performance across multiple external validation datasets.
- Zero-shot inference using CLIP provided accurate predictions for semantic features like nodule margin (AUROC: 0.812), nodule consistency (0.812), and pleural attachment (0.840).
Conclusions:
- The developed approach outperforms state-of-the-art models in lung cancer prediction across diverse clinical settings.
- The model provides explainable outputs, assisting clinicians in understanding prediction rationale and preventing model shortcuts.
- The method generalizes well across different clinical settings, offering a more reliable tool for lung cancer screening.

