Related Experiment Video
Updated: Jun 21, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Large Language Models as Tools to Generate Radiology Board-Style Multiple-Choice Questions
Neel P Mistry1, Huzaifa Saeed2, Sidra Rafique3
1College of Medicine, University of Saskatchewan, Saskatoon, Saskatchewan, Canada (N.P.M., H.S., H.O., S.J.A.); Department of Medical Imaging, Royal University Hospital, Saskatoon, Saskatchewan, Canada (N.P.M., H.O., S.J.A.).
Large language models (LLMs) show potential for creating radiology board-style multiple choice questions (MCQs). GPT-4 demonstrated high accuracy and quality comparable to existing exam materials, suggesting its utility for educators.
Area of Science:
- Artificial Intelligence in Medical Education
- Radiology Education Technology
Background:
- Radiology board-style multiple choice questions (MCQs) are crucial for resident assessment and learning.
- Developing high-quality MCQs requires significant educator time and expertise.
- Large language models (LLMs) offer a potential solution for efficient MCQ generation.
Purpose of the Study:
- To evaluate the capability of LLMs in generating radiology board-style MCQs, answers, and rationales.
- To compare the quality of LLM-generated MCQs with existing board examination items.
Main Methods:
- Two LLMs, Llama 2 and GPT-4, were used to create 104 MCQs based on the American Board of Radiology exam blueprint.
- Board-certified radiologists assessed LLM-generated MCQs and prior American College of Radiology (ACR) Diagnostic Radiology In-Training (DXIT) exam items on clarity, relevance, difficulty, distractors, and rationale adequacy.
- Radiologists were blinded to the source of the MCQs during the assessment.
Main Results:
- GPT-4 generated MCQs received scores comparable to ACR DXIT items across all assessed criteria, with near-perfect ratings.
- Llama 2-generated MCQs scored lower than ACR DXIT items but were still considered suitable.
- GPT-4 achieved 100% accuracy in answer generation, while Llama 2 achieved 69% accuracy.
Conclusions:
- Advanced LLMs like GPT-4 can effectively generate high-quality radiology board-style MCQs and rationales.
- LLMs can serve as valuable tools for radiology educators to enhance exam preparation materials and expand question banks.
- The use of LLMs may facilitate the broader application of MCQs as teaching and learning instruments in radiology.
More Related Videos
09:00Author Spotlight: Validation of SICOLE-R for Assessing Cognitive and Reading Skills in Spanish-Speaking Children and Its Role in Personalized Education
Published on: August 16, 2024
04:09Predicting Treatment Response to Image-Guided Therapies Using Machine Learning: An Example for Trans-Arterial Treatment of Hepatocellular Carcinoma
Published on: October 10, 2018
Related Concept Videos
Positron Emission Tomography
One of the main requirements of a PET scan is a positron-emitting radioisotope, which is produced in a cyclotron and then attached to a substance used by the part of the body...
Radiological Investigation II: MRI and Ventilation Perfusion Scan
Magnetic Resonance Imaging (MRI) and Ventilation Perfusion Scans are two radiological investigations that offer detailed diagnostic images of the body, particularly lung structures.
MRI
MRI uses magnetic fields and radiofrequency signals to distinguish between normal and abnormal tissues. This technology provides a more detailed diagnostic image than CT scans, enabling it to characterize pulmonary nodules, stage bronchogenic carcinoma, and evaluate inflammatory activity in...