Related Experiment Video
Updated: Jan 15, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
1.0K
Evaluating Hallucination and Diagnostic Reliability of LLMs on Medical Image-Based Multiple Choice Tasks
IEEE Journal of Biomedical and Health Informatics
|October 15, 2025
Summary
Evaluating large language models for medical diagnosis is crucial. While some models show accuracy, their reasoning often lacks clinical grounding, highlighting the need for explainability beyond correct answers.
Area of Science:
- Artificial Intelligence in Medicine
- Biomedical Informatics
- Medical Imaging Analysis
Background:
- Large language models (LLMs) are increasingly integrated into healthcare, offering potential for diagnostic support.
- Evaluating the reliability and clinical grounding of LLM reasoning is critical for safe application in sensitive decision-making.
- Current assessments often focus on diagnostic accuracy, neglecting the quality of the underlying reasoning process.
Purpose of the Study:
- To present a systematic framework for assessing both diagnostic correctness and explanation quality of LLMs in medical image-based tasks.
- To evaluate state-of-the-art LLMs on diverse medical cases across multiple specializations and imaging modalities.
- To introduce novel metrics for quantifying the reliability and clinical grounding of LLM-generated explanations.
Main Methods:
- Developed a framework to evaluate four LLMs on 30 medical cases across 10 specializations and 5 imaging modalities.
- Cases included diagnostic images with correct answers and distractors designed to probe reasoning vulnerabilities.
- Introduced metrics: hallucination rate, reasoning score, anatomical correctness, and grounding deviation score.
Main Results:
- Some LLMs achieved moderate diagnostic accuracy but often relied on superficial patterns or flawed logic.
- High grounding deviation scores were observed even for correct predictions, indicating a disconnect between answers and clinical reasoning.
- Weak or incorrect reasoning was the most frequent failure mode across evaluated models.
Conclusions:
- Focusing solely on diagnostic accuracy is insufficient; evaluating the 'how' and 'why' of LLM predictions is essential.
- The developed framework provides critical insights into LLM reasoning, promoting interpretability.
- This work supports the safer and more reliable integration of LLMs into biomedical diagnostic workflows.

