Related Experiment Video
Updated: Jan 26, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
On the transfer learning behavior of domain-specific vision-language models in screening mammography
Aisha Urooj Khan1, Gokul Ramasamy2, Muhammad Danish Khan3
1AI Innovation Hub, Mayo Clinic, Phoenix, AZ, USA.
None:
Vision-Language models have shown remarkable performance for natural images and text. Given the homology of the anatomy, high gray-scale image dimension, and the unbalanced datasets, the traditional VLMs do not adapt well to radiological applications. In this work, we empirically adapted image encoder trained within domain-specific VLMs to be applied in two downstream tasks for 2D mammogram image analysis: tissue density estimation and BI-RADS prediction. We study the transfer learning behavior using linear probing, fine-tuning, and online self distillation. We analyze that knowledge driven domain-specific VLM backbones with frozen weights perform better than MammoClip VLM model as well as supervised baselines such as ViT and CNNs even with only 5% of training data. Generalization capabilities are further studied of these models on two external datasets.
Related Concept Videos
Vision
Language
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
Color Vision
Components of Language
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...

