Related Experiment Video
Updated: Jan 17, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
1.0K
Enhancing 3D Medical Image Understanding With Pretraining Aided by 2D Multimodal Large Language Models
IEEE Journal of Biomedical and Health Informatics
|September 15, 2025
Summary
Med3DInsight enhances 3D medical image understanding by integrating 3D encoders with 2D multimodal large language models (MLLMs). This novel framework achieves state-of-the-art performance in segmentation and classification without human annotations.
Area of Science:
- Medical Imaging
- Artificial Intelligence
- Computer Vision
Background:
- Current 3D medical image analysis methods, including convolution and transformer-based self-supervised learning (SSL), often struggle with deep semantic understanding.
- Multimodal large language models (MLLMs) show potential for improving image comprehension through text integration.
Purpose of the Study:
- To introduce Med3DInsight, a novel pretraining framework designed to enhance 3D medical image understanding by leveraging 2D MLLMs.
- To develop a scalable multimodal representation learning approach for 3D medical data without requiring manual annotations.
Main Methods:
- Integration of 3D image encoders with 2D MLLMs using a plane-slice-aware transformer module.
- Employment of partial optimal transport-based alignment for increased noise tolerance in LLM-generated content.
- Development of a self-supervised learning framework for multimodal 3D medical representation learning.
Main Results:
- Achieved state-of-the-art performance on medical image segmentation and classification tasks across CT and MRI datasets.
- Outperformed existing self-supervised learning methods in downstream task evaluations.
- Demonstrated robustness to noise in LLM-generated content through optimal transport alignment.
Conclusions:
- Med3DInsight offers a new paradigm for multimodal 3D medical representation learning, improving semantic comprehension.
- The framework can be seamlessly integrated into existing 3D medical image analysis networks to boost performance.
- The approach enables scalable, annotation-free learning for 3D medical imaging.

