Frozen Large-Scale Pretrained Vision-Language Models are the Effective Foundational Backbone for Multimodal Breast

Summary

This study introduces a multimodal deep-learning model for breast cancer prediction, utilizing frozen vision-language models with mammogram data. The approach significantly improves prediction accuracy, especially for limited data scenarios, outperforming traditional methods.