Related Experiment Video
Updated: May 24, 2025

04:48
Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography
Published on: November 30, 2022
2.6K
Language Augmentation in CLIP for Improved Anatomy Detection on Multi-modal Medical Images
Summary
This study introduces an automated system for generating comprehensive medical image descriptions across the entire body using vision-language models. The novel approach significantly improves upon existing methods for multi-modal radiology report generation.
Area of Science:
- Medical Imaging
- Artificial Intelligence
- Computer Vision
Background:
- Vision-language models are increasingly used for multi-modal classification in medicine.
- Current automated radiology report generation is limited to specific modalities or body regions.
- A need exists for models capable of generating entire-body, multi-modal descriptions.
Purpose of the Study:
- To automate the generation of standardized body station(s) and organ lists for multi-modal MR and CT radiological images.
- To address the gap in current research for whole-body, multi-modal image description.
Main Methods:
- Leveraged Contrastive Language-Image Pre-training (CLIP) for model refinement.
- Conducted experiments including baseline model fine-tuning.
- Incorporated station(s) as a superset and applied image and language augmentations.
Main Results:
- Achieved a 47.6% performance improvement over the baseline PubMedCLIP model.
- Demonstrated successful automation of standardized body station(s) and organ lists across the whole body.
- Validated the effectiveness of augmentations and superset approach for improved correlation.
Conclusions:
- The proposed approach effectively automates comprehensive, multi-modal radiological image descriptions.
- This advancement has the potential to enhance radiology report generation and clinical workflow.
- Future work can build upon this model for more sophisticated medical image analysis.

