Comparing Diagnostic Accuracy of Radiologists versus GPT-4V and Gemini Pro Vision Using Image Inputs from Diagnosis

Pae Sun Suh1, Woo Hyun Shim1, Chong Hyun Suh1

  • 1From the Department of Radiology and Research Institute of Radiology, University of Ulsan College of Medicine, Asan Medical Center, Olympic-ro 33, Seoul 05505, Republic of Korea (P.S.S., W.H.S., C.H.S., H.J.E., K.J.P., J.C., P.H.K., H.J.P., Y.A., H.Y.P.); Department of Radiology and Research Institute of Radiological Science and Center for Clinical Imaging Data Science, Yonsei University College of Medicine, Seoul, Republic of Korea (P.S.S.); Department of Medical Science, University of Ulsan College of Medicine, Asan Medical Institute of Convergence Science and Technology, Seoul, Republic of Korea (W.H.S., H.H., C.R.P.); Medical Research Institute, Ganneung Asan Hospital, University of Ulsan College of Medicine, Gangneung, Republic of Korea (Y.C.); Department of Internal Medicine, University of Ulsan College of Medicine, Asan Medical Center, Seoul, Republic of Korea (C.Y.W.); and Department of Pulmonary and Critical Care Medicine, Gumdan Top Hospital, Incheon, Republic of Korea (H.P.).

Radiology
|July 9, 2024
PubMed
Summary

Large language models (LLMs) like GPT-4V and Gemini Pro Vision showed improved diagnostic accuracy with higher temperature settings when analyzing medical images. While radiologists still outperform LLMs, GPT-4V shows promise as a diagnostic support tool.