Related Experiment Video
Updated: Mar 6, 2026

Optimized Management of Endovascular Treatment for Acute Ischemic Stroke
Published on: January 18, 2018
Should we leave paediatric emergency triage to artificial intelligence? A comparison of ChatGPT 4o and Grok 3
Emre Aygun1, Aysenur Imdat1, Nazan Dalgic2
1Department of Pediatrics, Şişli Hamidiye Etfal Training and Research Hospital, University of Health Sciences, Istanbul, Türkiye.
Background:
The growing number of patients in paediatric emergency departments requires fast and precise triage assessments. The implementation of large language models faces obstacles due to their limited interpretability. We aimed to compare the performance of ChatGPT 4o and Grok 3 with that of nurses and physicians in paediatric emergency triage.
Methods:
This prospective observational study evaluated paediatric emergency patients presenting to our paediatric emergency department between March and April 2025. Demographic data, chronic disease status, presenting complaints, and vital signs were documented. Patients were triaged according to ESI criteria by nurses, paediatric specialists (gold standard), ChatGPT 4o, and Grok 3. Inter-rater agreement was analysed using Cohen's kappa**. Cochran's Q and McNemar's tests were used for paired comparisons.*.
Results:
A total of 1,505 paediatric emergency patients were included in the analysis. No ESI-1 cases were observed; therefore, critical patients were defined as ESI-2. Nurses achieved 53.1% (95% CI: 50.6-55.6) accuracy in triage assessments, while ChatGPT 4o achieved 76.1% (95% CI: 73.9-78.2) and Grok 3 achieved 47.0% (95% CI: 44.5-49.6) accuracy (Cochran's Q = 275.68, p < 0.001). ChatGPT 4o showed good agreement with physicians (κ = 0.69). For critical patient identification, sensitivity was 37.2% for nurses, 82.9% for ChatGPT 4o, and 97.7% for Grok 3; however, Grok 3 demonstrated substantial over-triage (36.3%) and low positive predictive value (37.2%). ChatGPT 4o achieved the lowest mean absolute ESI error (0.25 ± 0.45). Nurses' critical patient recognition improved from 28.3% to 59.5% (p < 0.01) for children with chronic illnesses.
Conclusion:
ChatGPT 4o achieved the most favourable balance of sensitivity and specificity. The superior performance of nurses in recognising critically ill patients with chronic diseases suggests that AI systems should augment nursing expertise rather than replace it.
Related Concept Videos
Acute Kidney Injury V: Interprofessional Care
Comparing the Survival Analysis of Two or More Groups
Guidelines and Strategies for Safe Computer Charting
Maintain Confidentiality and Security:
Methods of Documentation II: POMR

