Related Experiment Video
Updated: Aug 5, 2026

09:09
Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
Case-matched retrieval improves textual alignment of LLM-generated radiology impressions
Vera Sorin1, Jeremy D Collins1, Lewis D Hahn1
1Department of Radiology, Mayo Clinic College of Medicine and Science, Mayo Clinic, Rochester, Minnesota, United States of America.
Plos One
|July 31, 2026
Summary
Retrieval-augmented generation (RAG) improved Large Language Models (LLMs) drafting of radiology impressions. Dynamic, case-matched retrieval enhanced alignment with original reports, though human review remains essential.
Area of Science:
- Artificial Intelligence in Medicine
- Natural Language Processing
- Medical Imaging Informatics
Background:
- Radiology impressions are crucial for clinical decision-making.
- Large Language Models (LLMs) can produce generic text, impacting impression quality.
- Retrieval-augmented generation (RAG) offers context-aware prompting for LLMs.
Purpose of the Study:
- To evaluate the effectiveness of RAG in improving LLM-generated radiology impressions.
- To compare different few-shot prompting strategies for LLM impression generation.
- To assess the impact of retrieval methods on the stylistic alignment of generated impressions.
Main Methods:
- Retrospective analysis of 11,998 CT pulmonary angiography (CTPA) reports.
- Development of a retrieval bank from 11,399 reports for few-shot prompting.
- Comparison of zero-shot, fixed few-shot, and dynamic retrieval-based few-shot prompting using GPT-4o and LLaMA 3.1-70B.
- Evaluation using ROUGE and BERTScore F1 metrics at varying temperatures and k values.
Main Results:
- Dynamic retrieval-based few-shot prompting significantly outperformed other methods (p < 0.05).
- Optimal performance was achieved with temperature 0 and k=10, increasing ROUGE-1 F1 scores.
- GPT-4o and LLaMA 3.1-70B showed improved impression generation with dynamic RAG.
Conclusions:
- Dynamic, case-matched retrieval enhances the alignment of LLM-generated CTPA impressions with reference standards.
- Automated text-similarity metrics indicate improved impression quality.
- Radiologist verification is still necessary prior to clinical implementation.
