Related Experiment Video
Updated: Sep 19, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Automated Classification of Radiation Oncology Safety Events Using Large Language Models: A Novel Approach to
Qiongge Li1, Jian Liu2, Xing Li3
1Department of Radiation Oncology, Inova Schar Cancer Institute, Fairfax, Virginia; School of Medicine, University of Virginia, Charlottesville, Virginia.
Purpose:
This study aims to evaluate the feasibility of using a large language model (LLM) to automate the classification of radiation oncology safety events and to outline a framework for future reporting systems that leverage artificial intelligence (AI)-assisted workflows.
Methods And Materials:
Retrospective safety event data were extracted from our institutional reporting system and processed using Python scripts for deidentification and formatting. GPT-5 (OpenAI) was accessed via an application programming interface and refined through iterative prompt engineering (a natural language process that does not modify model parameters) to classify incidents across multiple dimensions, including failure mode, severity, treatment type, discoverer's role, exclusion criteria, and occurring/discovering workflow stages. The model's performance was validated by three independent, blinded expert reviewers, with inter-rater agreement quantified using Cohen's κ.
Results:
On a blinded 80-incident validation set, the model's classification fell within the experts' group-accepted answer in 67.9% of dimension-level comparisons on average (96.2% for the inclusion decision; 85.3% for treatment type and 76.6% for occurred workflow). Agreement between the model and individual reviewers (mean Cohen's κ = 0.31) was comparable to agreement among the reviewers themselves (mean κ = 0.42), indicating performance approaching that of an independent expert. Severity scoring showed the greatest variability for both the model and the human reviewers (model-reviewer κ = 0.14; inter-reviewer κ = 0.22), highlighting it as the primary area for improvement. The comparable model-expert and expert-expert agreement supported the application of the model to full data set analysis without further optimization.
Conclusions:
LLM-based classification demonstrates strong potential for automating retrospective safety event tagging and streamlining future reporting workflows. In a proposed future system, staff would only need to provide a detailed narrative description while the AI model performs classification and prompts for human validation when uncertainty arises. This hybrid approach can improve efficiency, consistency, and scalability in radiation oncology safety reporting. To our knowledge, this is the first study to apply an LLM for automated classification of radiation oncology safety events, introducing a proof-of-concept framework with potential for broader applicability for AI-assisted incident reporting and quality improvement.