Related Experiment Video
Updated: May 5, 2026

05:38
Interaction between Phonological and Semantic Processes in Visual Word Recognition using Electrophysiology
Published on: June 29, 2021
2.9K
Hallucination filtering in radiology vision-language models using discrete semantic entropy.
Patrick Wienholt1,2, Sophie Caselitz3,4, Robert Siepmann3,4
1Lab for Artificial Intelligence in Medicine, Department of Diagnostic and Interventional Radiology, University Hospital RWTH Aachen, Aachen, Germany. pwienholt@ukaachen.de.
European Radiology
|February 20, 2026
Summary
Discrete semantic entropy (DSE) effectively detects hallucination-prone questions for vision-language models (VLMs). This method significantly improves accuracy in radiologic image analysis, enhancing VLM reliability for clinical applications.
Area of Science:
- Artificial Intelligence
- Medical Imaging
- Natural Language Processing
Background:
- Black-box vision-language models (VLMs) are increasingly used for radiologic image analysis.
- Hallucinations, or inaccurate information generation, pose a significant challenge to VLM reliability in clinical settings.
- Current methods for detecting VLM hallucinations are limited, especially in black-box scenarios.
Purpose of the Study:
- To evaluate the efficacy of discrete semantic entropy (DSE) in identifying questions that may lead to hallucinations in VLMs.
- To determine if DSE-based filtering can improve the accuracy of VLMs in radiologic image-based visual question answering (VQA).
Main Methods:
- A retrospective study utilized two public datasets (VQA-Med 2019 and a diagnostic radiology dataset).
- GPT-4o and GPT-4.1 models answered questions, with DSE computed from semantic response clusters.
- Accuracy was assessed before and after filtering questions with DSE > 0.3.
Main Results:
- Baseline accuracy for GPT-4o was 51.7% and for GPT-4.1 was 54.8%.
- After DSE filtering (DSE > 0.3), accuracy increased to 76.3% for GPT-4o and 63.8% for GPT-4.1.
- Accuracy gains were statistically significant across datasets, even after Bonferroni correction.
Conclusions:
- DSE reliably detects hallucinations in black-box VLMs by quantifying semantic inconsistency.
- This method significantly enhances diagnostic answer accuracy in radiologic VQA.
- DSE provides a practical filtering strategy to improve the safety and trustworthiness of clinical VLM applications.

