Related Experiment Video
Updated: Sep 11, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Evaluating prompt and data perturbation sensitivity in large language models for radiology reports classification
Vera Sorin1, Jeremy D Collins1, Alex K Bratt1
1Department of Radiology, Mayo Clinic College of Medicine and Science, Mayo Clinic, Rochester, MN 55905, United States.
Large language models (LLMs) show high accuracy in classifying pulmonary embolism (PE) in radiology reports. Prompt design and data quality significantly impact LLM performance, necessitating careful validation for clinical use.
Area of Science:
- Artificial Intelligence in Medical Imaging
- Natural Language Processing in Healthcare
- Radiology Report Analysis
Background:
- Large language models (LLMs) are increasingly explored for healthcare applications, particularly in natural language processing tasks.
- Accurate classification of medical reports, such as those for pulmonary embolism (PE), is critical for patient care.
- Understanding the performance limitations of LLMs in high-stakes clinical scenarios is essential.
Purpose of the Study:
- To evaluate the performance of Google's Large Language Models (LLMs) in classifying pulmonary CT angiography radiology reports for the presence of pulmonary embolism (PE).
- To assess the impact of various conditions, including different prompt designs and data perturbations, on LLM classification accuracy.
- To determine the robustness and variability of LLM performance across different configurations and cloud environments.
Main Methods:
- Retrospective analysis of 11,999 pulmonary CT angiography radiology reports.
- Evaluation of three Google LLMs: Gemini-1.5-Pro, Gemini-1.5-Flash-001, and Gemini-1.5-Flash-002.
- Ground truth established by concordance between a computer vision-based PE detection (CVPED) algorithm and multiple LLM runs, with manual review for discrepancies.
- Analysis of prompt design, data perturbations, and geographic cloud region effects on performance metrics.
Main Results:
- Overall accuracy across LLMs ranged from 0.953 to 0.996.
- A modified prompt achieved a recall rate up to 0.997.
- Few-shot prompting improved recall (up to 0.99), while chain-of-thought prompting generally decreased performance.
- Gemini-1.5-Flash-002 showed the highest robustness against data perturbations; Gemini-1.5+-Pro had minimal geographic variability, while Flash models were stable.
Conclusions:
- LLMs demonstrate high efficacy in classifying radiology reports for pulmonary embolism.
- Performance is sensitive to prompt engineering and data quality, highlighting the need for rigorous validation.
- Systematic evaluation is crucial before deploying LLMs in critical clinical decision-making processes.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
07:15Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
Published on: August 16, 2020
Related Concept Videos
Sensitivity, Specificity, and Predicted Value
Sensitivity is the...
Improving Translational Accuracy
Survival Tree
Building a Survival Tree
Constructing a...