Related Experiment Video
Updated: Jul 28, 2026

07:13
Multimodal Cross-Device and Marker-Free Co-Registration of Preclinical Imaging Modalities
Published on: October 27, 2023
Multimodal Bidirectional Direct Preference Optimization and Instruction Fine-Tuning for Medical Image Understanding
IEEE Journal of Biomedical and Health Informatics
|June 25, 2026
Summary
This study introduces a two-stage framework to improve multimodal large language models (MLLMs) for medical imaging. The approach enhances accuracy and trustworthiness in generating radiology reports and medical images, reducing hallucinations.
Area of Science:
- Artificial Intelligence
- Medical Imaging
- Natural Language Processing
Background:
- Multimodal large language models (MLLMs) show promise but struggle with medical image nuances, particularly in radiology.
- Current supervised fine-tuning methods for MLLMs in healthcare often produce unreliable, hallucinated outputs, limiting clinical trust.
- The need for trustworthy AI in clinical decision support necessitates improved MLLM capabilities for medical data.
Purpose of the Study:
- To develop a novel two-stage fine-tuning framework to enhance the performance and reliability of MLLMs in medical imaging tasks.
- To address the limitations of existing methods by reducing hallucinations and improving the accuracy and clinical relevance of generated radiology reports.
- To improve the quality of medical image generation and establish MLLMs as trustworthy tools for multimodal clinical assistance.
Main Methods:
- A VQ-GAN-based visual tokenizer transforms medical images into discrete tokens, aligned with language tokens for visual-language instructional fine-tuning.
- The framework incorporates autoregressive tasks for both image and text generation, enabling interpretation of imaging-specific instructions across modalities.
- An improved Direct Preference Optimization (DPO) method is utilized, generating dispreferred data from distorted images to mitigate hallucinations effectively.
Main Results:
- The proposed framework significantly enhances the accuracy, faithfulness, and clinical relevance of generated radiology reports.
- Experimental results show substantial improvements in the quality of medical image generation.
- The method effectively mitigates hallucinations, a critical issue in applying MLLMs to clinical decision support.
Conclusions:
- The developed two-stage fine-tuning framework offers a reliable solution for multimodal clinical assistance using MLLMs.
- This approach advances trustworthy AI applications in healthcare by improving MLLM performance in medical diagnosis and reporting.
- The findings highlight the potential of MLLMs for precision diagnosis and reliable integration into clinical workflows.