Related Experiment Video
Updated: Jun 9, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Use of Large Language Models to Enhance Failure Mode and Effects Analysis: A Case Study
Saurabh S Nair1, Laurence Court1, Raphael Douglas1
1Department of Radiation Physics, The University of Texas MD Anderson Cancer Center, Houston, Texas, USA.
Purpose:
Failure Mode and Effects Analysis (FMEA) is widely used in radiation oncology to proactively identify and mitigate risks, but it is time-consuming and depends heavily on expert experience. This study evaluated whether large language models (LLMs) can supplement traditional expert-driven FMEA by identifying novel failure modes within the Radiation Planning Assistant (RPA) workflow.
Methods And Materials:
A multidisciplinary team of board-certified medical physicists, quality assurance engineers, and software developers independently used 4 LLMs (ChatGPT-4, Gemini 2.5 Pro, phi4-reasoning-14B, and OpenAI/oss-120B) to generate potential failure modes across the RPA contouring and planning workflow. Team members used diverse prompting strategies, including supplementary materials such as RPA user guides, as context. Each failure mode was first rated for severity, occurrence, and detectability by the LLMs, then independently rescored by experts using the TG-100 framework to enable comparison. The highest-risk modes, based on expert scoring, were subsequently reviewed with 2 clinical user groups in South Africa.
Results:
The 4 LLMs collectively generated 190 candidate failure modes. After review for relevance and duplication, 79 unique and interpretable modes were retained for analysis. Among these, 3 exceeded the 125 risk priority number threshold from a prior study, all related to staff accountability and role ambiguity. On average, LLMs assigned higher severity (7.3 vs 4.1), similar occurrence (2.8 vs 3.3), and lower detectability (5.4 vs 2.8) scores, producing higher mean RPNs (110 vs 36). Clinical users from 2 centers in South Africa confirmed that several artificial intelligence-identified risks were plausible, particularly those tied to workflow accountability.
Conclusions:
LLMs can broaden risk discovery in FMEA by surfacing contextually relevant and previously unrecognized failure modes. However, expert oversight remains essential for validating and prioritizing risks. Artificial intelligence should be viewed as a complementary tool that enhances, rather than replaces, human judgment in radiation therapy safety assessments.
Related Concept Videos
Typical Model Studies
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
Modeling and Similitude
Design Consideration
The factor of safety is another key aspect...
Design Example: Creating a Hydraulic Model of a Dam Spillway