Related Experiment Video
Updated: Jun 5, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Comparing Large Language Model and Human Reader Accuracy with New England Journal of Medicine Image Challenge Case
Pae Sun Suh1, Woo Hyun Shim1, Chong Hyun Suh1
1From the Department of Radiology, Research Institute of Radiological Science and Center for Clinical Imaging Data Science, Yonsei University College of Medicine, Seoul, Republic of Korea (P.S.S.); Department of Radiology and Research Institute of Radiology (W.H.S., C.H.S., K.J.P., P.H.K., S.J.C., Y.A., S.P., H.Y.P., N.E.O.), Department of Medical Science, Asan Medical Institute of Convergence Science and Technology (W.H.S., H.H.), and Department of Internal Medicine (C.Y.W.), Asan Medical Center, University of Ulsan College of Medicine, Olympic-ro 33, Songpa-gu, 05505 Seoul, Republic of Korea; University of Ulsan College of Medicine, Seoul, Republic of Korea (M.W.H.); Department of Orthopaedic Surgery, Seoul Seonam Hospital, Republic of Korea (S.T.C.); and Department of Pulmonary and Critical Care Medicine, Gumdan Top Hospital, Incheon, Republic of Korea (H.P.).
Large language models (LLMs) show promise in interpreting radiologic images, outperforming medical students but not experienced radiologists. LLM accuracy improves with longer text inputs, regardless of image use.
Area of Science:
- Artificial Intelligence in Medicine
- Medical Imaging Analysis
- Natural Language Processing
Background:
- Multimodal large language models (LLMs) integrating text and vision are increasingly used, yet their diagnostic accuracy in radiology remains under scrutiny.
- The study addresses the need to evaluate LLM performance against human expertise in interpreting complex medical cases.
Purpose of the Study:
- To assess the accuracy of leading LLMs in interpreting radiologic images compared to human readers of varying experience levels.
- To identify factors influencing LLM performance, specifically the impact of image versus text inputs and text length.
Main Methods:
- Retrospective review of 272 New England Journal of Medicine Image Challenge cases from 2005 to 2024.
- Evaluation of four LLMs (GPT-4V, GPT-4o, Gemini 1.5 Pro, Claude 3) and a panel of human readers (radiologists, clinicians, medical student).
- Subgroup analysis examined LLM accuracy with/without image inputs and with short versus long text descriptions, using multivariable logistic regression and generalized estimating equations for statistical comparison.
Main Results:
- GPT-4o achieved the highest LLM accuracy (59.6%), surpassing a medical student (47.1%) but not junior faculty radiologists (80.9%) or an in-training radiologist (70.2%).
- LLM accuracy was consistent with or without image inputs (59.6% vs 54.0%) and significantly improved with longer text descriptions (odds ratio range, 3.2–6.6).
- Human reader accuracy was not affected by text length, highlighting a key difference in performance characteristics.
Conclusions:
- LLMs demonstrate significant accuracy in interpreting radiologic cases, offering a valuable tool that can outperform less experienced human readers.
- LLM performance is notably influenced by the length of textual information provided, with longer descriptions yielding better results.
- While promising, LLMs still lag behind experienced radiologists, and their reliance on text length suggests areas for future development and optimization in multimodal medical AI.
More Related Videos
04:09Predicting Treatment Response to Image-Guided Therapies Using Machine Learning: An Example for Trans-Arterial Treatment of Hepatocellular Carcinoma
Published on: October 10, 2018
05:56Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application
Published on: April 14, 2023