Related Experiment Video
Updated: Apr 20, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Performance of Large Language Models on Diagnostic Radiology Board-Style Questions: A Comparative Evaluation of
Randall Aziz1, Sydney Stewart1, Rebecca Liscomb1
1USF Health Morsani College of Medicine, University of South Florida, Tampa.
Objective:
The objective of this study was to compare the diagnostic accuracy and internal consistency of GPT-4o (Generative Pre-Trained Transformer-4 omni), Perplexity AI (artificial intelligence), and OpenEvidence when applied to text-based, specialty-level radiology board questions.
Methods:
A total of 161 text-based multiple-choice questions from the American College of Radiology (ACR) Diagnostic Radiology In-Training Examination were administered across three independent runs for each large language model (LLM). Questions containing images were excluded. All three models were accessed through their respective public Web interfaces. A final answer was assigned to each model based on majority vote across the three runs (two out of three). If all three responses differed, the third (last) response was selected. Our selected answer was then compared with the ACR reference key. Internal consistency as well as agreement between each model's final answer and the ACR reference key was assessed using Cohen's kappa. In addition, descriptive statistics were used to analyze performance by radiology subspecialty. SPSS version 30 was used for all statistical analyses, and P<0.05 were considered statistically significant.
Results:
Perplexity AI demonstrated the highest agreement with the ACR reference key (κ=0.883, P<0.001), followed by OpenEvidence (κ=0.858, P<0.001), and GPT-4o (κ=0.709, P<0.001). All models showed high internal consistency; however OpenEvidence was the only LLM to demonstrate absolute internal consistency (κ=1.00 for all three runs). Perplexity AI showed the least variability across the 14 radiology subspecialties.
Conclusion:
Emerging LLMs such as Perplexity AI and OpenEvidence may offer greater diagnostic reliability than general-purpose models in radiology-specific contexts.
Related Concept Videos
Radiological Investigation II: MRI and Ventilation Perfusion Scan
Magnetic Resonance Imaging (MRI) and Ventilation Perfusion Scans are two radiological investigations that offer detailed diagnostic images of the body, particularly lung structures.
MRI
MRI uses magnetic fields and radiofrequency signals to distinguish between normal and abnormal tissues. This technology provides a more detailed diagnostic image than CT scans, enabling it to characterize pulmonary nodules, stage bronchogenic carcinoma, and evaluate inflammatory activity in...
Positron Emission Tomography
One of the main requirements of a PET scan is a positron-emitting radioisotope, which is produced in a cyclotron and then attached to a substance used by the part of the body...
Radiological Investigation III: Pulmonary Angiogram and PET Scan
Pulmonary Angiogram
A Pulmonary Angiogram is an invasive procedure involving injecting a contrast medium through a catheter threaded into the pulmonary artery or the right side of the heart to visualize the pulmonary vasculature. Computed Tomography (CT) scans have mainly replaced this...
Radiological Investigation I: X-ray and CT
Imaging Studies II: Positron Emission Tomography and Scintigraphy
Fundamental Principles of PET
