Related Experiment Video
Updated: Jan 7, 2026

11:27
A Modified Sonographic Algorithm for Image Acquisition in Life-Threatening Emergencies in the Critically Ill Newborn
Published on: April 7, 2023
7.2K
Clinical decision-making in neonatal gastrointestinal surgical emergencies: comparison between ChatGPT and human
Seyoon Kim1, Dongho Choi1, Joonhyuk Son1
1Department of Surgery, Hanyang University College of Medicine, Seongdong-gu, The Republic of Korea.
World Journal of Pediatric Surgery
|December 29, 2025
Summary
GPT-4o shows promise in aiding clinical decisions for neonatal gastrointestinal surgical emergencies (NGSEs). It outperformed general surgery residents and attendings, nearing pediatric surgery expert levels in key areas.
Area of Science:
- Medical Informatics
- Artificial Intelligence in Medicine
- Pediatric Surgery
Background:
- Neonatal gastrointestinal surgical emergencies (NGSEs) present complex diagnostic and management challenges.
- Rapid clinical decision-making is crucial to minimize morbidity and mortality in neonates with NGSEs.
- The study explores the utility of advanced AI, specifically GPT-4o, in supporting these critical decisions.
Purpose of the Study:
- To assess the efficacy of GPT-4o as a clinical decision-support tool for neonatal gastrointestinal surgical emergencies (NGSEs).
- To compare GPT-4o's performance against general surgery (GS) residents, GS attendings, and pediatric surgery (PS) attendings in managing NGSE cases.
Main Methods:
- Five challenging NGSE cases were transformed into structured questions with clinical and radiological data.
- GPT-4o underwent 10 iterations per case, with performance scored against evaluations by 10 GS residents, 10 GS attendings, and 10 PS attendings.
- Statistical analysis compared GPT-4o's scores with those of the human expert groups.
Main Results:
- GPT-4o achieved a mean score of 89.9%, significantly outperforming GS residents (p<0.001) and GS attendings (p<0.001).
- GPT-4o's performance was comparable to PS attendings (89.9% vs. 95.4%, p=0.021), particularly in management, final diagnosis, and surgical planning.
- GPT-4o scored lower than PS attendings in differential diagnosis (87.8% vs. 92.8%) and diagnostic plan (75.0% vs. 93.8%).
Conclusions:
- GPT-4o demonstrates significant potential as a supplementary decision-support tool for neonatal gastrointestinal surgical emergencies (NGSEs).
- Its performance rivals that of experienced pediatric surgeons in critical management aspects.
- Further validation in real-world clinical settings is necessary before widespread adoption for NGSEs.

