Related Experiment Video
Updated: May 27, 2026

04:56
Laryngeal Mask Airway (LMA) Placement in a Neonatal Patient Simulator Using a Non-Inflatable Supraglottic Airway (SGA)
Published on: July 14, 2023
Feasibility and challenges of large language model (LLM)-generated neonatal resuscitation simulations: a multicenter
Chenguang Xu1,2, Ming Zhou3, Dianna Wang4
1NICU, The University of Hong Kong-Shenzhen Hospital, Shenzhen, China.
Summary
Large language models (LLMs) show promise for generating neonatal resuscitation simulation scenarios. ChatGPT scenarios were feasible and comparable to traditional methods, though instructor supervision is crucial for effective implementation.
Area of Science:
- Medical Education
- Artificial Intelligence in Healthcare
- Simulation Technology
Background:
- Simulation-based training (SBT) enhances neonatal resuscitation outcomes but demands significant instructor resources.
- Large language models (LLMs) offer potential for dynamic, contextual scenario generation in medical training.
- Gaps exist in understanding the feasibility and challenges of LLM-generated scenarios for neonatal resuscitation.
Purpose of the Study:
- To evaluate the feasibility and challenges of using LLM-generated simulation scenarios in neonatal resuscitation training.
- To compare LLM-generated scenarios with established methods (NRP®, RETAIN) in terms of design and fidelity.
- To assess AI hallucination and qualitative aspects of LLM-generated scenarios.
Main Methods:
- A prospective, multicenter pilot study involving 16 simulation scenarios (4 each from ChatGPT-4o, DeepSeek-R1, NRP®, RETAIN).
- Nine blinded instructors evaluated scenarios using a modified Jeffries Simulation Design Scale (JSDS).
- Comparison included AI hallucination, qualitative feedback, and adherence to the Neonatal Resuscitation Program® (NRP®) algorithm.
Main Results:
- ChatGPT-generated scenarios showed comparable overall evaluation scores to NRP® scenarios.
- ChatGPT scenarios scored higher than NRP® in debriefing design.
- DeepSeek and RETAIN scenarios received lower scores for overall evaluation and fidelity; DeepSeek also showed information provision issues and NRP® algorithm violations.
Conclusions:
- LLM-generated simulation scenarios, particularly from ChatGPT, may be feasible for neonatal resuscitation SBT under instructor supervision.
- LLM scenarios require careful evaluation for NRP® deviations and gaps before implementation.
- Further research on educational outcomes and learner feedback is essential for integrating LLM-generated simulations.
