Related Experiment Video
Updated: May 22, 2025

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
474
Generative Large Language Models Trained for Detecting Errors in Radiology Reports
Cong Sun1, Kurt Teichman2, Yiliang Zhou1
1Department of Population Health Sciences, Weill Cornell Medicine, 575 Lexington Ave, New York, NY 10022.
Radiology
|May 20, 2025
Summary
Large language models (LLMs) significantly improve medical proofreading by detecting errors in radiology reports. Fine-tuned LLMs, like Llama-3, demonstrated high accuracy in identifying negation, left/right, interval, and transcription errors.
Area of Science:
- Artificial Intelligence in Medicine
- Medical Informatics
- Radiology Reporting
Background:
- Large language models (LLMs) show potential for medical proofreading, but their use in radiology report error detection is limited.
- Developing AI tools for accurate radiology report analysis is crucial for patient safety and clinical decision-making.
Purpose of the Study:
- To develop and evaluate generative LLMs for detecting various errors in radiology reports.
- To assess the performance of different LLMs and prompting strategies in medical proofreading tasks.
Main Methods:
- A dataset of synthetic and real-world radiology reports (MIMIC-CXR) was created, including error-free and erroneous reports.
- Errors were categorized into negation, left/right, interval change, and transcription types.
- Models including Llama-3, GPT-4, and BiomedBERT were fine-tuned using zero-shot, few-shot, and fine-tuning strategies, with performance evaluated by F1 scores and radiologist review.
Main Results:
- The fine-tuned Llama-3-70B-Instruct model achieved the highest overall F1 score of 0.780, with specific high scores for transcription errors (0.828).
- Radiologist review confirmed the model's ability to detect errors, with 99 reports confirmed by both reviewers and 163 by at least one reviewer.
Conclusions:
- Generative LLMs, when fine-tuned on diverse radiology report datasets, substantially enhance the accuracy of medical proofreading.
- These AI models offer a promising solution for improving the quality and reliability of radiology reports.
Related Concept Videos
Leaky Scanning
5.0K
During most eukaryotic translation processes, the small 40S ribosome subunit scans an mRNA from its 5' end until it encounters the first start AUG codon. The large 60S ribosomal subunit then joins the smaller one to initiate protein synthesis. The location of the translation initiation is largely determined by the nucleotides near the start codon as there may be multiple translation initiation sites present on the mRNA. Marilyn Kozak discovered that the sequence RCCAUGG (where R...
5.0K
Positron Emission Tomography
3.9K
Positron emission tomography (PET) is a medical imaging technique involving radiopharmaceuticals — substances that emit short-lived radiation. Although the first PET scanner was introduced in 1961, it took 15 more years before radiopharmaceuticals were combined with the technique and revolutionized its potential.
One of the main requirements of a PET scan is a positron-emitting radioisotope, which is produced in a cyclotron and then attached to a substance used by the part of the body...
One of the main requirements of a PET scan is a positron-emitting radioisotope, which is produced in a cyclotron and then attached to a substance used by the part of the body...
3.9K
Improving Translational Accuracy
8.5K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
8.5K
Nonsense-mediated mRNA Decay
10.4K
The Upf proteins that carry out nonsense-mediated decay (NMD) are found in all eukaryotic organisms, including humans. Each protein has an individual role, but they need to work in collaboration. Upf1 is an ATP-dependent RNA helicase that unwinds the RNA helix. Because Upf1 can unwind any RNA, Upf2 and Upf3 are required to help Upf1 discriminate between nonsense and normal mRNAs.
Usually, Upf3 binds to an Exon Junction Complex (EJC) at mRNA splice sites. If a ribosome fully translates the mRNA,...
Usually, Upf3 binds to an Exon Junction Complex (EJC) at mRNA splice sites. If a ribosome fully translates the mRNA,...
10.4K

