Evaluating Hallucination and Diagnostic Reliability of LLMs on Medical Image-Based Multiple Choice Tasks

Summary

Evaluating large language models for medical diagnosis is crucial. While some models show accuracy, their reasoning often lacks clinical grounding, highlighting the need for explainability beyond correct answers.