Related Experiment Video
Updated: Jun 21, 2025

Detection of Architectural Distortion in Prior Mammograms via Analysis of Oriented Patterns
Published on: August 30, 2013
Comparing Diagnostic Accuracy of Radiologists versus GPT-4V and Gemini Pro Vision Using Image Inputs from Diagnosis
Pae Sun Suh1, Woo Hyun Shim1, Chong Hyun Suh1
1From the Department of Radiology and Research Institute of Radiology, University of Ulsan College of Medicine, Asan Medical Center, Olympic-ro 33, Seoul 05505, Republic of Korea (P.S.S., W.H.S., C.H.S., H.J.E., K.J.P., J.C., P.H.K., H.J.P., Y.A., H.Y.P.); Department of Radiology and Research Institute of Radiological Science and Center for Clinical Imaging Data Science, Yonsei University College of Medicine, Seoul, Republic of Korea (P.S.S.); Department of Medical Science, University of Ulsan College of Medicine, Asan Medical Institute of Convergence Science and Technology, Seoul, Republic of Korea (W.H.S., H.H., C.R.P.); Medical Research Institute, Ganneung Asan Hospital, University of Ulsan College of Medicine, Gangneung, Republic of Korea (Y.C.); Department of Internal Medicine, University of Ulsan College of Medicine, Asan Medical Center, Seoul, Republic of Korea (C.Y.W.); and Department of Pulmonary and Critical Care Medicine, Gumdan Top Hospital, Incheon, Republic of Korea (H.P.).
Large language models (LLMs) like GPT-4V and Gemini Pro Vision showed improved diagnostic accuracy with higher temperature settings when analyzing medical images. While radiologists still outperform LLMs, GPT-4V shows promise as a diagnostic support tool.
Area of Science:
- Artificial Intelligence in Medicine
- Medical Imaging Diagnostics
- Natural Language Processing
Background:
- The diagnostic capabilities of multimodal large language models (LLMs) with direct image input are not well understood.
- The influence of LLM temperature parameters on diagnostic accuracy requires further investigation.
Purpose of the Study:
- To evaluate GPT-4V and Gemini Pro Vision's ability to generate differential diagnoses from medical images.
- To compare LLM diagnostic performance at varying temperature settings (0, 0.5, 1) against expert radiologists.
Main Methods:
- Retrospective analysis of 190 Radiology Diagnosis Please cases (2008-2023).
- LLMs received images, patient history, and figure legends; generated 3 differential diagnoses, repeated 5 times at each temperature.
- Radiologists' diagnoses were compared to LLM outputs; accuracy assessed across models, temperatures, and subspecialties.
Main Results:
- Overall accuracy increased with temperature for both GPT-4V (41%-49%) and Gemini Pro Vision (29%-39%).
- Radiologists achieved higher overall accuracy (61%) than Gemini Pro Vision at T1 (39%, P < .001).
- GPT-4V at T1 (49%) showed no statistically significant difference compared to radiologists (P = .02), but radiologists outperformed LLMs in most subspecialties.
Conclusions:
- Increasing temperature settings improved diagnostic accuracy for GPT-4V and Gemini Pro Vision using direct image inputs.
- GPT-4V demonstrated promising potential as a supportive diagnostic tool, despite slightly lower performance than radiologists.
- LLMs show potential to aid in diagnostic decision-making, warranting further research into their clinical application.
More Related Videos
14:08Automated Midline Shift and Intracranial Pressure Estimation based on Brain CT Images
Published on: April 13, 2013
04:09Predicting Treatment Response to Image-Guided Therapies Using Machine Learning: An Example for Trans-Arterial Treatment of Hepatocellular Carcinoma
Published on: October 10, 2018