Related Experiment Video
Updated: Jan 29, 2026

Eye Tracking During Visually Situated Language Comprehension: Flexibility and Limitations in Uncovering Visual Context Effects
Published on: November 30, 2018
Visual Large Language Models in Radiology: A Systematic Multimodel Evaluation of Diagnostic Accuracy and
Marc Sebastian von der Stück1, Roman Vuskov1, Simon Westfechtel1
1Department of Diagnostic and Interventional Radiology, University Hospital RWTH Aachen, 52074 Aachen, Germany.
None:
Visual large language models (VLLMs) are discussed as potential tools for assisting radiologists in image interpretation, yet their clinical value remains unclear. This study provides a systematic and comprehensive comparison of general-purpose and biomedical VLLMs in radiology. We evaluated 180 representative clinical images with validated reference diagnoses (radiography, CT, MRI; 60 each) using seven VLLMs (ChatGPT-4o, Gemini 2.0, Claude Sonnet 3.7, Perplexity AI, Google Vision AI, LLaVA-1.6, LLaVA-Med-v1.5). Each model interpreted the image without and with clinical context. Mixed-effects logistic regression models assessed the influence of model, modality, and context on diagnostic performance and hallucinations (fabricated findings or misidentifications). Diagnostic accuracy varied significantly across all dimensions (p ≤ 0.001), ranging from 8.1% to 29.2% across models, with Gemini 2.0 performing best and LLaVA performing weakest. CT achieved the best overall accuracy (20.7%), followed by radiography (17.3%) and MRI (13.9%). Clinical context improved accuracy from 10.6% to 24.0% (p < 0.001) but shifted the model to rely more on textual information. Hallucinations were frequent (74.4% overall) and model-dependent (51.7-82.8% across models; p ≤ 0.004). Current VLLMs remain diagnostically unreliable, heavily context-biased, and prone to generating false findings, which limits their clinical suitability. Domain-specific training and rigorous validation are required before clinical integration can be considered.
More Related Videos
Related Concept Videos
Uncertainty in Measurement: Accuracy and Precision
Language
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...
Improving Translational Accuracy
Heart Failure IV: Classification and Diagnostic Evaluation
Positive Symptoms Schizophrenia: Hallucinations and Delusions
Hallucinations
Hallucinations in...
Irritable Bowel Syndrome II: Clinical Features and Diagnostic Evaluation
Irritable Bowel Syndrome (IBS) is classified into subtypes based on the predominant bowel habits as determined by the Bristol Stool Form Scale (BSFS). The subtypes are:

