Related Experiment Video
Updated: May 26, 2025

Radiation Planning Assistant - A Streamlined, Fully Automated Radiotherapy Treatment Planning System
Published on: April 11, 2018
Differentiating between GPT-generated and human-written feedback for radiology residents
Zier Zhou1, Arsalan Rizwan2, Nick Rogoza3
1School of Medicine, Queen's University, Kingston, ON, Canada.
Large language models (LLMs) like GPT-3.5 struggle to generate specific, nuanced feedback for radiology residents compared to human experts. Current LLMs also perform poorly at distinguishing between human-written and AI-generated educational assessments.
Area of Science:
- Medical Education
- Artificial Intelligence in Healthcare
- Radiology Training
Background:
- Competency-based medical education (CBME) in Canadian radiology programs necessitates increased faculty assessment.
- The proliferation of large language models (LLMs) prompts an evaluation of their utility in generating narrative feedback for resident assessments.
Purpose of the Study:
- To compare human-written feedback with GPT-3.5-generated feedback for radiology residents.
- To assess the ability of human raters and GPT-3.5 to differentiate between human and AI-generated feedback.
Main Methods:
- 110 resident feedback comments from a Canadian Diagnostic Radiology program were analyzed.
- 110 synthetic comments were generated using GPT-3.5 and mixed with actual human comments.
- Two faculty raters and GPT-3.5 attempted to identify the source of each comment.
Main Results:
- Human feedback was generally longer and more specific than GPT-3.5 generated comments.
- Differentiation between human and AI feedback was harder when comments were vague.
- Human raters achieved 80.5% accuracy in source identification, while GPT-3.5 achieved only 50%.
Conclusions:
- GPT-3.5 currently falls short of human expert capabilities in providing specific, nuanced feedback for radiology residents.
- GPT-3.5 demonstrates a lower capacity than humans in distinguishing between authentic and synthetic feedback.
- Findings may inform the development of advanced AI algorithms for improved feedback generation and faculty development support.
More Related Videos
09:03Radiotracer Administration for High Temporal Resolution Positron Emission Tomography of the Human Brain: Application to FDG-fPET
Published on: October 22, 2019
03:14Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Related Concept Videos
X-ray Imaging
Methods of Documentation III: PIE
Positron Emission Tomography
One of the main requirements of a PET scan is a positron-emitting radioisotope, which is produced in a cyclotron and then attached to a substance used by the part of the body...
Methods of Documentation II: POMR
Data Reporting and Recording
Guidelines for Writing Outcome
Patient outcomes reflect the patient's response to the goal rather than what the nurse aims to achieve. Terminology should be observable and measurable to avoid the reader's interpretation. The desired outcome should be realistic and achievable in the designated care timeframe. Expected outcomes should align with adjunctive therapies. The outcome should enhance care...