Related Experiment Video
Updated: Aug 5, 2026

04:16
Fine-Tuning Large Language Models Using Entity Hallucination Index for Text Summarization
Published on: January 9, 2026
MedHallTune: An Instruction-Tuning Benchmark for Mitigating Medical Hallucination in Vision-Language Models
IEEE Transactions on Medical Imaging
|August 3, 2026
Summary
This study introduces MedHallTune, a benchmark to address hallucinations in medical vision-language models (VLMs). Fine-tuning with this dataset improves VLM reliability for clinical applications.
Area of Science:
- Artificial Intelligence
- Medical Informatics
- Computer Vision
Background:
- Vision-language models (VLMs) are increasingly used in healthcare.
- Model hallucinations, generating incorrect yet plausible outputs, pose risks to clinical decision-making, diagnosis, and treatment.
- Existing VLMs require robust evaluation and mitigation strategies for medical applications.
Purpose of the Study:
- To introduce MedHallTune, a large-scale benchmark for evaluating and mitigating hallucinations in medical VLMs.
- To provide a comprehensive resource for assessing VLM performance in clinical contexts.
- To enhance the trustworthiness and reliability of VLMs in healthcare.
Main Methods:
- Developed MedHallTune, a benchmark with over 100,000 images and 1,000,000 instruction pairs, including hallucination and non-hallucination samples.
- Constructed the dataset using GPT-based generation and filtering, with manual verification of the evaluation split by medical professionals.
- Evaluated current medical and general VLMs on MedHallTune using metrics like clinical accuracy, relevance, detail level, and risk level.
Main Results:
- Fine-tuning existing VLMs with MedHallTune significantly improved their ability to manage hallucinations.
- The benchmark enhanced the zero-shot performance of models on downstream visual-question-answering tasks.
- Models fine-tuned with MedHallTune demonstrated increased reliability for practical medical applications.
Conclusions:
- MedHallTune is an effective benchmark for evaluating and mitigating hallucinations in medical VLMs.
- The proposed fine-tuning approach enhances VLM trustworthiness and clinical applicability.
- Public availability of the benchmark, codes, and prompts will foster further research in reliable medical AI.