Related Experiment Video
Updated: Jan 15, 2026

A Multimodal Imaging Framework to Advance Phenotyping of Living Label-free Breast Cancer Cells
Published on: August 22, 2025
Leveraging pretrained vision-language model for enhanced breast cancer diagnosis with multi-view mammography
Xuxin Chen1, Yuheng Li1,2, Mingzhe Hu1,3
1Department of Radiation Oncology and Winship Cancer Institute, Emory University, Atlanta, Georgia, USA.
Background:
Although fusion of information from multiple views of mammograms plays an important role to increase accuracy of breast cancer detection, developing multi-view mammograms-based computer-aided diagnosis (CAD) schemes still faces big challenges and no such CAD schemes have been used in clinical practice.
Purpose:
To overcome these challenges, we investigate a new approach based on the concept of contrastive language-image pre-training (CLIP), which has sparked interest across various medical imaging tasks. The aim is to solve the challenges in: (1) effectively adapting the single-view CLIP for multi-view feature fusion and (2) efficiently fine-tuning this parameter-dense model with limited samples and computational resources.
Methods:
We introduce a unique Mammo-CLIP, the first multi-modal framework to process multi-view mammograms and corresponding simple texts. Mammo-CLIP uses an early feature fusion strategy to learn multi-view relationships in four mammograms acquired from the craniocaudal (CC) and mediolateral oblique (MLO) views of the left and right breasts. To enhance learning efficiency, plug-and-play adapters are added into CLIP's image and text encoders for fine-tuning the model efficiently and limiting updates to about 1% of the parameters. For framework evaluation, we assembled two datasets retrospectively. The first dataset, comprising 470 malignant and 479 benign cases, was used for few-shot fine-tuning and internal evaluation of the proposed Mammo-CLIP via 5-fold cross-validation. The second dataset, including 60 malignant and 294 benign cases, was used to test generalizability of Mammo-CLIP.
Results:
Mammo-CLIP outperforms the state-of-the-art (SOTA) cross-view transformer evaluated using areas under ROC curves (AUC = 0.841 ± 0.017 vs. 0.817 ± 0.012 and 0.837 ± 0.034 vs. 0.807 ± 0.036) on both datasets. It also surpasses previous two CLIP-based methods by 20.3% and 14.3% in AUC.
Conclusions:
The proposed Mammo-CLIP demonstrates superior breast cancer diagnosis performance compared to SOTA methods. This study highlights the potential of applying the finetuned vision-language models for developing multi-view, image-text-based CAD schemes of breast cancer.
