Related Experiment Video
Updated: Jun 12, 2026

Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography
Published on: November 30, 2022
Decoupling Visual Parsing and Diagnostic Reasoning for Vision-Language Models (GPT-4o and GPT-5): Analysis Using
Dae Hee Han1, Eui Jin Hwang2, Soon Ho Yoon2,3
1Department of Radiology, Seoul St. Mary's Hospital, College of Medicine, The Catholic University of Korea, Seoul, Korea.
None:
BACKGROUND. Vision-language models (VLMs) have potential to identify findings on radiologic imaging (i.e., visual parsing) and translate findings into diagnoses (i.e., diagnostic reasoning). Current VLMs have shown insufficient performance to support clinical integration. OBJECTIVE. The purpose of our study was to evaluate the separate contributions of visual parsing and diagnostic reasoning toward GPT-based VLMs' performance in generating correct diagnoses for thoracic imaging. METHODS. This retrospective study included 128 publicly available thoracic imaging cases from the Korean Society of Thoracic Imaging quiz platform (accessed on June 15, 2025). Two VLMs (GPT-4o and GPT-5) processed cases, separately when patient metadata and images were inputted and when patient metadata and radiologist-generated image descriptions were inputted. The models provided five ranked differential diagnoses for each case; when metadata and images were inputted, the models first provided a summary of imaging findings. The proportion of cases for which the models' five differential diagnoses included the correct diagnosis was determined (i.e., top-5 accuracy). The performance of quiz participants, who interpreted cases using metadata and images, was extracted from the platform. The quality of the model-provided image summaries was scored on a 4-point scale (4 = best score). Logistic regression analyses assessed associations between model image summary scores and diagnostic performance. Diagnostic concordance was assessed between models' top-ranked diagnoses and quiz participants' top-10 differential diagnoses. RESULTS. Top-5 accuracy for GPT-4o and GPT-5 when metadata and images were inputted was 15.9% and 24.7% and when metadata and descriptions were inputted was 40.1% and 59.1%, respectively; quiz participants' pooled top-5 accuracy was 45.8%. Median image summary score was 2 for both models; these scores showed significant independent associations with a top-5 match (GPT-4o: OR = 5.95; GPT-5: OR = 2.77; p < .001). Concordance between models' top-ranked diagnosis and quiz participants' differential lists for GPT-4o and GPT-5 when metadata and images were inputted was 31.6% and 39.3% and when metadata and descriptions were inputted was 78.8% and 79.4%, respectively. CONCLUSION. Two VLMs showed limited ability to visually identify thoracic imaging findings but performed more favorably in generating accurate diagnoses when provided radiologist-generated descriptions. CLINICAL IMPACT. The results underscore the need for radiologist expertise in thoracic imaging interpretation and identify visual image parsing rather than diagnostic reasoning as the principal limitation constraining VLM performance.
More Related Videos
07:11Functional Magnetic Resonance Imaging (fMRI) of the Visual Cortex with Wide-View Retinotopic Stimulation
Published on: December 8, 2023
02:09Multi-modal Pulmonary Imaging: Using Complementary Information from CT and Hyperpolarized 129Xe MRI to Evaluate Lung Structure-Function
Published on: April 12, 2024
Related Concept Videos
Computed Tomography
The technique was invented in the 1970s and is based on the principle that as X-rays pass through the body, they are absorbed or reflected at different levels. In the technique, a patient lies on a motorized platform while a computerized axial tomography (CAT) scanner rotates...
Positron Emission Tomography
One of the main requirements of a PET scan is a positron-emitting radioisotope, which is produced in a cyclotron and then attached to a substance used by the part of the body being...
Radiological Investigation II: MRI and Ventilation Perfusion Scan
Magnetic Resonance Imaging (MRI) and Ventilation Perfusion Scans are two radiological investigations that offer detailed diagnostic images of the body, particularly lung structures.
MRI
MRI uses magnetic fields and radiofrequency signals to distinguish between normal and abnormal tissues. This technology provides a more detailed diagnostic image than CT scans, enabling it to characterize pulmonary nodules, stage bronchogenic carcinoma, and evaluate inflammatory activity in...
Radiological Investigation III: Pulmonary Angiogram and PET Scan
Pulmonary Angiogram
A Pulmonary Angiogram is an invasive procedure involving injecting a contrast medium through a catheter threaded into the pulmonary artery or the right side of the heart to visualize the pulmonary vasculature. Computed Tomography (CT) scans have mainly replaced this...
Imaging Studies III: Computed Tomography