Related Experiment Video
Updated: May 18, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Automated Analysis of Radiation Oncology Incident Reports Using Large Language Models: A Multi-Institutional
Nathan Dobranski1, Natalie N Viscariello2, Garrett M Pitcher3
1Department of Physics and Astronomy, Louisiana State University, Baton Rouge, Louisiana.
Purpose:
Patient safety incident reporting in radiation oncology requires expert analysis that is time-intensive and subject to variability. This study evaluates the technical feasibility of large language model (LLM) automation for incident report analysis across multiple cancer centers.
Methods And Materials:
A locally deployed LLM system was developed for automated incident report summarization and taxonomy assignment. The system processed Radiation Oncology Incident Learning System reports from 2 anonymized cancer centers with 600 total expert evaluations. Round 1 established baseline performance with 495 evaluations (institution 1, n = 257; institution 2, n = 238) using Mistral 7B and Mixtral 8x7B models. Round 2 incorporated enhanced prompt engineering, retrieval-augmented generation, and advanced models (Gemma3 27B, DeepSeek R1 70B, Llama3.1 70B) with 105 evaluations (institution 1, n = 60; institution 2, n = 45). Clinical experts evaluated outputs using 5-point ordinal scales (1 = nonsensical to 5 = excellent), with mixed-effects statistical analysis.
Results:
Multirater analysis demonstrated statistically significant performance improvements between rounds across both institutions. Institution 1 achieved substantial improvements: summary scores (3.34 ± 1.17 → 4.20 ± 0.84; d = 0.78; P < .001) and tag scores (3.28 ± 0.98 → 4.32 ± 0.70; d = 1.11; P < .001). Institution 2 showed improvements: summary scores (3.61 ± 1.24 → 4.02 ± 0.97; d = 0.34; P = .055) and tag scores (3.79 ± 1.10 → 4.40 ± 0.81; d = 0.58; P < .001). Round 2 high-performance thresholds (≥4): institution 1 achieved 80.0% for summaries, and 86.7% for tags; institution 2 achieved 68.9% and 84.4%, respectively. Processing time increased 3-fold (5.6 → 19.2 seconds per report) with substantial performance gains.
Conclusions:
This study shows feasibility of automated radiation oncology incident analysis using locally deployed LLMs, with effect sizes ranging from small to large (d = 0.34-1.11) across institutions. Performance improvements establish technical feasibility for further development, while institutional variation suggests that site-specific factors may influence system effectiveness, which warrants further investigation.
