Related Experiment Video
Updated: Jun 8, 2025

06:21
Diffuse Optical Spectroscopy for the Quantitative Assessment of Acute Ionizing Radiation Induced Skin Toxicity Using a Mouse Model
Published on: May 27, 2016
8.1K
Zero-Shot Medical Phrase Grounding With Off-the-Shelf Diffusion Models
IEEE Journal of Biomedical and Health Informatics
|November 8, 2024
Summary
This study introduces a novel zero-shot phrase grounding method for medical image localization using a Latent Diffusion Model. The approach effectively utilizes textual reports to pinpoint pathologies in scans without additional training.
Area of Science:
- Medical Imaging
- Artificial Intelligence
- Computer Vision
Background:
- Accurate localization of pathological regions in medical scans typically requires extensive bounding box annotations.
- Readily available free-text reports offer a weaker, yet accessible, form of supervision for localization tasks.
Purpose of the Study:
- To perform phrase grounding (localization using textual guidance) in medical images in a zero-shot manner.
- To leverage a Foundation Model (Latent Diffusion Model) for medical image localization without model fine-tuning.
Main Methods:
- Utilized a Latent Diffusion Model, exploiting its cross-attention mechanisms for implicit visual-textual feature alignment.
- Developed feature selection and post-processing strategies for refinement without learnable parameters.
- Evaluated the zero-shot approach against state-of-the-art contrastive learning methods.
Main Results:
- The proposed method demonstrated competitive performance against state-of-the-art approaches on a chest X-ray benchmark.
- Achieved superior average performance in mean IoU and AUC-ROC metrics compared to existing methods.
- Showcased the efficacy of zero-shot learning for medical image localization.
Conclusions:
- Zero-shot phrase grounding using Foundation Models is a viable and effective strategy for medical image localization.
- The Latent Diffusion Model's inherent properties facilitate alignment for localization tasks.
- This approach reduces the need for extensive manual annotations, making localization more accessible.

