Hallucination-Aware Multimodal Benchmark for Gastrointestinal Image Analysis with Large Vision-Language Models

Bidur Khanal1, Sandesh Pokhrel2, Sanjay Bhandari2

  • 1Rochester Institute of Technology, Rochester, NY, USA.

Medical Image Computing and Computer-Assisted Intervention : MICCAI ... International Conference on Medical Image Computing and Computer-Assisted Intervention
|March 25, 2026
PubMed
Summary

Vision-Language Models (VLMs) in medicine can hallucinate, generating inaccurate reports. A new dataset and hallucination-aware finetuning method improve VLM accuracy for gastrointestinal image analysis.