Related Experiment Video
Updated: Jul 5, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Generative Large Language Models for Detection of Speech Recognition Errors in Radiology Reports.
Reuben A Schmidt1, Jarrel C Y Seah1, Ke Cao1
1From the Department of Medical Imaging, Western Health, Footscray, Australia (R.A.S., L.L., W.L.); Alfred Health, Harrison.ai, Monash University, Clayton, Australia (J.C.Y.S.); Department of Surgery, Western Precinct, University of Melbourne, Melbourne, Australia (K.C., J.Y.); and Department of Surgery, Western Health, Melbourne, Australia (J.Y.).
Generative large language models (LLMs) can detect speech recognition errors in radiology reports. GPT-4 showed high accuracy, indicating potential for automated error detection in medical imaging reports.
Area of Science:
- Medical Informatics
- Artificial Intelligence
- Radiology
Background:
- Speech recognition errors in radiology reports can impact patient care.
- Manual error detection is time-consuming and prone to human error.
Purpose of the Study:
- To evaluate the efficacy of generative large language models (LLMs) in detecting speech recognition errors within radiology reports.
- To compare the performance of various LLMs, including GPT-4, in identifying these errors.
Main Methods:
- A dataset of 3233 CT and MRI reports was analyzed for speech recognition errors by radiologists.
- Five generative LLMs (GPT-3.5-turbo, GPT-4, text-davinci-003, Llama-v2-70B-chat, Bard) were tested for error detection.
- Prompt engineering was utilized to optimize LLM performance, with manual detection serving as the reference standard.
Main Results:
- GPT-4 achieved high accuracy, with F1 scores of 86.9% for clinically significant errors and 94.3% for non-clinically significant errors.
- Other LLMs showed varied performance, with Bard exhibiting the lowest accuracy.
- Factors such as longer reports and resident dictation were associated with increased error rates.
Conclusions:
- Advanced generative LLMs, particularly GPT-4, demonstrate significant potential for the automated detection of speech recognition errors in radiology reports.
- This technology could enhance the accuracy and efficiency of radiology reporting quality assurance.
- Further research may explore integrating LLMs into clinical workflows for real-time error flagging.
More Related Videos
Related Concept Videos
Leaky Scanning
Nonsense-mediated mRNA Decay
Usually, Upf3 binds to an Exon Junction Complex (EJC) at mRNA splice sites. If a ribosome fully translates the mRNA,...
Positron Emission Tomography
One of the main requirements of a PET scan is a positron-emitting radioisotope, which is produced in a cyclotron and then attached to a substance used by the part of the body...

