Related Experiment Video
Updated: Jan 7, 2026

Introduction of an Integrated Pathology Image Management, Artificial Intelligence, and Reporting System
Published on: July 11, 2025
Evaluating Generative Artificial Intelligence as an Educational Tool for Radiology Resident Report Drafting
Antonio Verdone1, Aidan Cardall2, Fardeen Siddiqui2
1Department of Radiology, New York University Grossman School of Medicine, New York, New York.
Objective:
Radiology residents require timely, personalized feedback to develop accurate image analysis and reporting skills. Increasing clinical workload often limits attendings' ability to provide guidance. This study evaluates a HIPAA-compliant Generative Pretrained Transformer (GPT)-4o system that delivers automated feedback on breast imaging reports drafted by residents in real clinical settings.
Methods:
We analyzed 5,000 resident-attending report pairs from routine practice at a multisite US health system. GPT-4o was prompted with clinical instructions to identify common errors and provide feedback. A reader study using 100 report pairs was conducted. Four attending radiologists and four residents independently reviewed each pair, determined whether predefined error types were present, and rated GPT-4o's feedback as helpful or not. Agreement between GPT and readers was assessed using percent match. Interreader reliability was measured with Krippendorff's α. Educational value was measured as the proportion of cases rated helpful.
Results:
Three common error types were identified: (1) omission or addition of key findings, (2) incorrect use or omission of technical descriptors, and (3) final assessment inconsistent with findings. GPT-4o showed strong agreement with attending consensus: 90.5%, 78.3%, and 90.4% (Cohen's κ: 0.790, 0.550, and 0.615) across error types. Interreader reliability among all eight readers showed moderate to substantial variability (α = 0.767, 0.595, 0.567). When each reader was individually replaced with GPT-4o and interreader agreement among seven readers and GPT was recalculated, the effect was not statistically significant (Δ = -0.004 to 0.002, all P > .05). GPT's feedback was rated helpful in most cases: 89.8%, 83.0%, and 92.0%.
Discussion:
ChatGPT-4o can reliably identify key educational errors. It may serve as a scalable tool to support radiology education.
Related Concept Videos
Non-equilibrium in the Cell
Radiological Investigation I: X-ray and CT
Radiological Investigation II: MRI and Ventilation Perfusion Scan
Magnetic Resonance Imaging (MRI) and Ventilation Perfusion Scans are two radiological investigations that offer detailed diagnostic images of the body, particularly lung structures.
MRI
MRI uses magnetic fields and radiofrequency signals to distinguish between normal and abnormal tissues. This technology provides a more detailed diagnostic image than CT scans, enabling it to characterize pulmonary nodules, stage bronchogenic carcinoma, and evaluate inflammatory activity in...