Related Experiment Video
Updated: Jun 6, 2025

Hybrid µCT-FMT imaging and image analysis
Published on: June 4, 2015
Toward Foundation Models in Radiology? Quantitative Assessment of GPT-4V's Multimodal and Multianatomic Region
Quirin D Strotzer1, Felix Nieberle1, Laura S Kupke1
1From the Institute of Radiology (Q.D.S., L.S.K., G.N., A.K.M., S.M., I.E., J.R., C.W., C.S., O.W.H., A.S.) and Department of Cranio- and Maxillofacial Surgery (F.N.), University of Regensburg Medical Center, Franz-Josef-Strauss-Allee 11, 93053 Regensburg, Germany; Department of Radiology, Division of Neuroradiology, Massachusetts General Hospital, Harvard Medical School, Boston, Mass (Q.D.S.); Department of Radiology, Bayreuth Medical Center, Bayreuth, Germany (M.S.); Center of Neuroradiology, medbo District Hospital and University Medical Center Regensburg, Regensburg, Germany (I.W., C.W.); and Department of Radiology, Donaustauf Hospital, Donaustauf, Germany (O.W.H.).
GPT-4V can identify medical image types and locations but struggles with detecting and classifying abnormalities. This large vision-language model showed a high false-positive rate in interpreting radiologic images.
Area of Science:
- Artificial Intelligence
- Medical Imaging
- Radiology
Background:
- Large language models (LLMs) show promise in processing medical text.
- GPT-4V, a vision-language model, has potential in medical imaging analysis.
- Quantitative assessment of GPT-4V for radiologic image interpretation is needed.
Purpose of the Study:
- To quantitatively evaluate GPT-4V's performance in interpreting unseen radiologic images.
- To assess the accuracy, sensitivity, and specificity of GPT-4V's diagnostic reports.
- To compare GPT-4V's performance against human radiologists.
Main Methods:
- Retrospective analysis of single abnormal and healthy control images across neuroradiology, cardiothoracic, and musculoskeletal radiology.
- GPT-4V interpretation via API for free-text reports and binary classification tasks.
- Performance metrics (accuracy, sensitivity, specificity) compared to a first-year resident and four board-certified radiologists.
Main Results:
- GPT-4V achieved 100% accuracy in identifying imaging modality and 99.2% in anatomic region.
- Diagnostic accuracy for free-text reports varied widely (0% for pneumothorax to 90% for brain tumor).
- GPT-4V demonstrated high false-positive rates (86.5% for free-text, 67.7% for binary classification) and lower sensitivity/specificity compared to human readers.
Conclusions:
- The initial version of GPT-4V can recognize medical image content, modality, and anatomy.
- GPT-4V currently fails to reliably detect, classify, or rule out abnormalities in radiologic images.
- Further development is required for GPT-4V to be a reliable tool in diagnostic radiology.
More Related Videos
07:13Author Spotlight: An Efficient and Robust Software for Automated Fusion of Multiple Preclinical Imaging Modalities
Published on: October 27, 2023
09:55Radiosynthesis, Quality Control, and Small Animal Positron Emission Tomography Imaging of 68Ga-Labelled Nano Molecules
Published on: October 4, 2024
Related Concept Videos
Computed Tomography
The technique was invented in the 1970s and is based on the principle that as X-rays pass through the body, they are absorbed or reflected at different levels. In the technique, a patient lies on a motorized platform while a computerized axial tomography (CAT) scanner rotates...
Imaging Studies II: Positron Emission Tomography and Scintigraphy
Fundamental Principles of PET
Positron Emission Tomography
One of the main requirements of a PET scan is a positron-emitting radioisotope, which is produced in a cyclotron and then attached to a substance used by the part of the body...