Can a Generative Artificial Intelligence Model Be Used to Create Mass Casualty Incident Simulation Scenarios? A
Sergio M Navarro1, Angie G Atkinson2, Ege Donagay3
1Division of Trauma, Critical Care, and General Surgery, Department of Surgery, Mayo Clinic, 200 1st St. SW, Rochester, MN 55905, USA.
Introduction:
Mass casualty incident (MCI) simulation scenarios are developed based on detailed review and planning by multidisciplinary trauma teams. This study aimed to assess the feasibility of using generative artificial intelligence (AI) in developing mass casualty trauma simulation scenarios. The study evaluated a range of mass casualty trauma simulation scenarios generated from a public generative artificial intelligence platform based on publicly available data with a validated objective simulation scoring tool.
Methods:
Using a large language model (LLM) platform (ChatGPT4, OpenAI, San Francisco, CA, USA), 10 complex MCI trauma simulation scenarios were generated based on publicly available US reported trauma data. Each scenario was evaluated by two Advanced Trauma Life Support (ATLS) certified raters based on the Simulation Scenario Evaluation Tool (SSET), a validated scoring tool out of 100 points. The tool scoring is based on learning objectives, tasks for performance, clinical progression, debriefing criteria, and resources. Two publicly available mass casualty trauma scenarios were similarly evaluated as controls. Revision and recommended feedback was provided for the scenarios, with review time recorded. Post-revision scenarios were evaluated. Interrater reliability was calculated based on Intraclass Correlation Coefficients (2, k) (ICCs). For the scenarios, scores and review times were reported as medians with interquartile range (IQR) as 25th and 75th percentiles.
Results:
Ten mass casualty trauma simulation scenarios were generated by an LLM, producing a total of 62 simulated patients. The initial LLM-generated scenarios demonstrated a median SSET score of 78.5 (IQR 74-82), substantially lower than the median score of 94 (IQR 93-95) observed in publicly available scenarios. The interrater reliability ICC for the LLM-generated scenarios was 0.965 and 1.00 for publicly available scenarios. Following secondary human revision and iterative refinement, the LLM-generated scenarios improved, achieving a median SSET score of 94 (IQR 93-96) with an interrater reliability ICC of 0.7425.
Conclusions:
The feasibility study suggests that a structured, collaborative workflow combining LLM-based generation with expert human review may enable a new approach to mass casualty trauma simulation scenario creation. LLMs hold promise as a scalable tool for the development of MCI training materials. However, consistent human oversight, quality assurance processes, and governance frameworks remain essential to ensure clinical accuracy, safety, and educational value.
Related Concept Videos
Steps in Outbreak Investigation
Non-equilibrium in the Cell

