Related Experiment Video
Updated: Sep 27, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Augmenting patient safety surveillance in radiation oncology with large language model-based root cause analysis: A
Yuntao Wang1, Mariluz De Ornelas1, Matthew T Studenski1
1Department of Radiation Oncology, University of Miami, Miami, Florida, United States of America.
Abstract:
We investigated the potential utility of large language models (LLMs) in supporting patient safety efforts. Specifically, we evaluated the reasoning capabilities of LLMs in performing root cause analysis (RCA) of radiation oncology incidents using narrative reports from the Radiation Oncology Incident Learning System (RO-ILS). We prompted four state-of-the-art LLMs, Gemini 2.5 Pro, GPT-4o, o3, and Grok 3, with the "Background and Incident Overview" sections from 19 publicly available RO-ILS cases. Each model was instructed to perform RCA and generate root causes, lessons learned, and suggested actions using a standardized prompt based on AAPM RCA guidelines. Model outputs were evaluated using a combination of objective semantic similarity metrics (cosine similarity via Sentence Transformer), semi-subjective assessments (precision, recall, F1-score, expert-adjudicated PPV (Positive Predictive Value), hallucination rate and performance criteria including relevance, comprehensiveness, quality of justification and quality of solution), and subjective ratings (reasoning quality and overall performance) by five board-certified medical physicists. LLMs demonstrated satisfactory performance across evaluation metrics. All models exhibited some degree of hallucination, ranging from 11% to 61%. All the evaluated LLMs demonstrated comparable baseline capabilities in objective causal extraction, and Gemini 2.5 Pro exhibited the highest overall performance score among 4 models. Statistically significant differences were observed among models in expert-adjudicated PPV, hallucination rate, and subjective ratings (p < 0.05). LLMs delivered promising results as assistive tools for RCA in radiation oncology, with the ability to generate relevant and accurate analyses aligned with expert expectations. LLMs may support incident analysis and contribute to quality improvement efforts to advance patient safety in clinical radiation oncology practice.