Related Experiment Video
Updated: Jun 7, 2026

07:45
Orthotopic Transplantation of Breast Tumors as Preclinical Models for Breast Cancer
Published on: May 18, 2020
5.9K
Frozen Large-Scale Pretrained Vision-Language Models are the Effective Foundational Backbone for Multimodal Breast
IEEE Journal of Biomedical and Health Informatics
|March 3, 2025
Summary
This study introduces a multimodal deep-learning model for breast cancer prediction, utilizing frozen vision-language models with mammogram data. The approach significantly improves prediction accuracy, especially for limited data scenarios, outperforming traditional methods.
Area of Science:
- Medical Imaging Analysis
- Artificial Intelligence in Healthcare
- Oncology
Background:
- Breast cancer is a major global health issue for women.
- Multimodal data integration from Picture Archiving and Communication Systems (PACS) and Electronic Health Records (EHRs) shows potential for improved breast cancer prediction.
- Traditional image-tabular models face challenges with non-aligned clinical data.
Purpose of the Study:
- To develop and evaluate a multimodal deep-learning model for breast cancer prediction using mammograms.
- To assess the performance of integrating frozen large-scale pretrained vision-language models with clinical data.
- To compare the proposed model against traditional methods and investigate its efficacy in limited data scenarios.
Main Methods:
- A multimodal deep-learning model was developed, integrating frozen pretrained vision-language models with mammogram datasets.
- The model utilized a lightweight trainable classifier alongside the frozen vision-language components.
- Performance was evaluated on two public breast cancer datasets (CBIS-DDSM and EMBED), including scenarios with limited data (BI-RADS 3 cases).
Main Results:
- The multimodal model demonstrated superior performance and stability compared to traditional image-tabular models.
- Significant improvements in Area Under the Curve (AUC) were observed: CBIS-DDSM validation (0.867 to 0.902) and test (0.803 to 0.830); EMBED validation (0.780 to 0.805).
- In limited data scenarios (BI-RADS 3), AUC increased from 0.91 to 0.96 (CBIS-DDSM test) and 0.79 to 0.83 (validation set).
Conclusions:
- Frozen large-scale pretrained vision-language models are effective for multimodal breast cancer prediction.
- The proposed approach offers superior performance and stability over conventional methods, particularly with diverse, multi-institutional healthcare data.
- This research highlights the potential of vision-language models to enhance breast cancer prediction accuracy and address data integration challenges.
Related Concept Videos
Mouse Models of Cancer Study
Mice have long served as models for studying human biology and pathology because of their phylogenetic and physiological similarity with humans. They are also easy to maintain and breed in the laboratory, and hence, many inbred strains are now available for research. Studies on mice have contributed immeasurably to our understanding of cancer biology.
The development of transgenic, knockout, and knock-in mice has led to an exponential increase in their use as model organisms in research,...
The development of transgenic, knockout, and knock-in mice has led to an exponential increase in their use as model organisms in research,...
Mouse Models of Cancer Study
Mice have long served as models for studying human biology and pathology because of their phylogenetic and physiological similarity with humans. They are also easy to maintain and breed in the laboratory, and hence, many inbred strains are now available for research. Studies on mice have contributed immeasurably to our understanding of cancer biology.
The development of transgenic, knockout, and knock-in mice has led to an exponential increase in their use as model organisms in research,...
The development of transgenic, knockout, and knock-in mice has led to an exponential increase in their use as model organisms in research,...

