Related Experiment Video
Updated: Jun 30, 2026

09:41
A Pipeline for 3D Multimodality Image Integration and Computer-assisted Planning in Epilepsy Surgery
Published on: May 20, 2016
12.8K
Read like a radiologist: Efficient vision-language model for 3D medical imaging interpretation
Changsun Lee1, Sangjoon Park2, Cheong-Il Shin3
1Kim Jaechul Graduate School of AI, Korea Advanced Institute of Science and Technology (KAIST), Daejeon, Republic of Korea.
Medical Image Analysis
|April 16, 2026
Summary
A new model, MS-VLM, enhances 3D medical image interpretation by mimicking radiologists. It generates more coherent and clinically relevant radiology reports from 3D medical imaging data.
Area of Science:
- Artificial Intelligence
- Medical Imaging
- Computer Vision
Background:
- Medical vision-language models (VLMs) show promise in 2D image interpretation but struggle with 3D medical imaging due to computational demands and data limitations.
- Existing 3D VLMs often use sub-volumetric features, leading to correlated representations and neglecting crucial slice-specific details in 3D medical images.
Purpose of the Study:
- To introduce MS-VLM, a novel model designed to overcome the limitations of current 3D medical vision-language models.
- To develop a VLM that mimics the human radiologist's workflow for 3D medical image interpretation, capturing inter-slice dependencies effectively.
Main Methods:
- MS-VLM utilizes self-supervised 2D transformer encoders to learn volumetric representations from sequences of slice-specific features.
- The model processes 3D medical images without sub-volumetric patchification, allowing flexibility with varying slice lengths and multiple imaging planes/phases.
Main Results:
- MS-VLM demonstrated superior performance in radiology report generation on chest CT and rectal MRI datasets.
- The model produced more coherent and clinically relevant reports compared to existing methods.
- MS-VLM effectively captures inter-slice dependencies, improving volumetric representation from 3D medical images.
Conclusions:
- MS-VLM represents a significant advancement in 3D medical image interpretation.
- The model's ability to mimic radiologists' workflows enhances the robustness and clinical relevance of medical VLMs for 3D imaging.

