Related Experiment Video
Updated: Jul 28, 2026

Multimodal Cross-Device and Marker-Free Co-Registration of Preclinical Imaging Modalities
Published on: October 27, 2023
Multimodal Bidirectional Direct Preference Optimization and Instruction Fine-Tuning for Medical Image Understanding
None:
Although multimodal large language models (MLLMs) are advancing rapidly in general vision-language tasks, their ability to capture the subtle nuances of medical images, especially in radiology, is limited. Current methods primarily apply supervised fine-tuning, which often leads to hallucinated results and renders them untrustworthy in clinical decision support. As a way of overcoming this drawback, we suggest a two step finetuning framework. The first stage uses a VQ-GAN-based visual tokenizer takes medical images and transforms them into discrete tokens, and then aligned with the language token format. Both image and text generation are considered autoregressive tasks, based on text or image inputs. This step conducts a visual-language instructional fine tuning, which allows the model to interpret and follow imaging specific instructions in a variety of imaging modalities. In the second stage, we present an improved Direct Preference Optimization (DPO) method. We deliberately distort images to induce halluci nations and generate dispreferred data, while genuine data serve as the preferred reference. This refined DPO strategy effectively mitigates hallucinations. Experimental results demonstrate that the proposed framework significantly improves the accuracy, faithfulness, and clinical relevance of generated radiology reports, while also enhancing the quality of medical image generation. These findings high light the potential of MLLMs for reliable multimodal clinical assistance, supporting precision diagnosis and advancing trustworthy AI applications in healthcare.