Related Experiment Video
Updated: May 24, 2025

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
475
Leveraging a Vision-Language Model with Natural Text Supervision for MRI Retrieval, Captioning, Classification, and
Biorxiv : the Preprint Server for Biology
|March 3, 2025
Summary
This study introduces a novel framework for brain MRI analysis using natural language supervision, enabling versatile tasks like retrieval and question answering for Alzheimer's disease research.
Area of Science:
- Artificial Intelligence
- Medical Imaging
- Neuroscience
Background:
- Large multimodal models are widely used but raise concerns about data quality, domain relevance, and privacy in medical applications.
- Current deep learning models in radiology are often task-specific and lack natural language interaction capabilities.
Purpose of the Study:
- To develop a versatile framework for learning visual brain MRI concepts using natural language supervision.
- To enable multiple downstream tasks including MRI retrieval, captioning, classification, and visual question answering.
- To investigate the identification of factors affecting Alzheimer's disease (AD) in brain MRIs.
Main Methods:
- Utilized vector retrieval and contrastive learning for efficient concept learning.
- Employed self-supervised learning to pre-train separate text and image encoders.
- Jointly fine-tuned encoders to create a shared embedding space for cross-modal learning.
- Developed a retrieval and re-ranking mechanism with a transformer decoder for visual question answering.
Main Results:
- The framework successfully learns to identify factors influencing Alzheimer's disease (AD) through joint embedding and natural language supervision.
- Demonstrated the model's capability to perform multiple tasks: MRI retrieval, captioning, and classification.
- Showcased versatility through a retrieval/re-ranking mechanism and transformer decoder for visual question answering.
Conclusions:
- The proposed framework offers a general and versatile tool for radiologic research by integrating medical imaging with text.
- Enables diagnostic and prognostic assessments in AD research and assists clinicians by detecting radiologic features described in text.
- Represents a novel approach to radiologic research, enhancing the utility of multimodal models in healthcare.

