Related Experiment Videos
Do Accuracy Gains Reflect Genuine Visual Understanding? A Multi-Model Evaluation of Vision-Language Models in
Jiyoung Song1, Won Gi Jeong2, Hongseok Ko1
1Department of Radiology, Seoul National University Hospital, Seoul National University College of Medicine, 101 Daehak-Ro, Jongno-Gu, Seoul, 03080, Korea.
Abstract:
Diagnostic accuracy of vision-language models (VLMs) on radiology benchmarks is rising rapidly, but whether higher multiple-choice accuracy reflects genuine visual understanding, or merely the selection of a plausible option from a closed set, remains unclear. We retrospectively evaluated 13 VLMs, comprising paired latest and preceding versions from five proprietary families (GPT, Gemini, Claude, Grok, Mistral), one standalone, and two open-weights models, on 100 challenging thoracic cases (chest radiograph, CT, MR, PET) reformatted as single-best-answer multiple-choice questions, under image-input (primary) and description-input (non-visual) conditions, against 10 readers including three thoracic radiologists. For the top-performing family, all rationales accompanying correct and incorrect answers were graded for reasoning errors (hallucination, missing critical finding, wrong detail) by two thoracic radiologists. Under image-input, Gemini 3.0 Pro achieved the highest Top-1 accuracy (57.0%), close in aggregate to the thoracic-reader mean (57.3%), and four of five latest proprietary models gained 16.0 to 19.6 points over their predecessors (all p < 0.01). However, only 55.7% of Gemini 3.0 Pro's rationales were free of errors: among correct answers, 80.7% were soundly reasoned, leaving nearly one in five correct diagnoses resting on hallucinated, missing, or incorrect findings; among incorrect answers, only 22.5% reflected sound reasoning, and 36.4% contained hallucinations. Across all 13 models, stated confidence systematically exceeded observed accuracy (ECE, 18.2 points), with the gap widening at higher confidence. Description-input raised accuracy further (Gemini 3.0 Pro, 79.7%) but relied on diagnosis-aware descriptions unavailable in practice. Even at the highest accuracy observed here, answer-level performance can overstate visual understanding, pointing to the need for more appropriate evaluation methods that assess whether a model's reasoning is faithful to the image, not only whether its answer is correct.