Related Experiment Video
Updated: Apr 20, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Performance of Large Language Models on Diagnostic Radiology Board-Style Questions: A Comparative Evaluation of
Randall Aziz1, Sydney Stewart1, Rebecca Liscomb1
1USF Health Morsani College of Medicine, University of South Florida, Tampa.
Perplexity AI and OpenEvidence show higher diagnostic accuracy on radiology board questions than GPT-4o. These emerging AI models may be more reliable for specialized medical contexts.
Area of Science:
- Artificial Intelligence in Medicine
- Medical Education Technology
- Radiology Informatics
Background:
- Large language models (LLMs) are increasingly used in various fields, including medicine.
- Evaluating the diagnostic accuracy and internal consistency of LLMs for specialized medical examinations is crucial.
- Radiology board questions require high diagnostic precision and consistency.
Purpose of the Study:
- To compare the diagnostic accuracy and internal consistency of GPT-4o, Perplexity AI, and OpenEvidence on text-based radiology board questions.
- To assess the performance of these LLMs across different radiology subspecialties.
- To determine if specialized AI models offer improved reliability over general-purpose models in a medical context.
Main Methods:
- 161 text-based multiple-choice questions from the American College of Radiology (ACR) Diagnostic Radiology In-Training Examination were used.
- Each LLM (GPT-4o, Perplexity AI, OpenEvidence) was run three times independently.
- Diagnostic agreement was assessed using Cohen's kappa, with performance analyzed by subspecialty.
Main Results:
- Perplexity AI achieved the highest agreement with the ACR reference key (κ=0.883), followed by OpenEvidence (κ=0.858) and GPT-4o (κ=0.709).
- All models demonstrated high internal consistency, with OpenEvidence showing absolute consistency (κ=1.00).
- Perplexity AI exhibited the least performance variability across 14 radiology subspecialties.
Conclusions:
- Emerging LLMs like Perplexity AI and OpenEvidence show promising diagnostic reliability for radiology-specific applications.
- Specialized AI models may outperform general-purpose LLMs in accuracy and consistency for medical board examinations.
- Further research into AI applications in medical education and diagnostics is warranted.
Related Concept Videos
Radiological Investigation II: MRI and Ventilation Perfusion Scan
Magnetic Resonance Imaging (MRI) and Ventilation Perfusion Scans are two radiological investigations that offer detailed diagnostic images of the body, particularly lung structures.
MRI
MRI uses magnetic fields and radiofrequency signals to distinguish between normal and abnormal tissues. This technology provides a more detailed diagnostic image than CT scans, enabling it to characterize pulmonary nodules, stage bronchogenic carcinoma, and evaluate inflammatory activity in...
Positron Emission Tomography
One of the main requirements of a PET scan is a positron-emitting radioisotope, which is produced in a cyclotron and then attached to a substance used by the part of the body...
Radiological Investigation III: Pulmonary Angiogram and PET Scan
Pulmonary Angiogram
A Pulmonary Angiogram is an invasive procedure involving injecting a contrast medium through a catheter threaded into the pulmonary artery or the right side of the heart to visualize the pulmonary vasculature. Computed Tomography (CT) scans have mainly replaced this...
Radiological Investigation I: X-ray and CT
Imaging Studies II: Positron Emission Tomography and Scintigraphy
Fundamental Principles of PET
