Related Experiment Video
Updated: Jun 3, 2025

07:31
Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
Published on: May 15, 2020
7.0K
Evaluation of Generative Artificial Intelligence Models in Predicting Pediatric Emergency Severity Index Levels
Brandon Ho, Meng Lu1, Xuan Wang1
1Department of Computer Science, Virginia Tech, Falls Church, VA.
Pediatric Emergency Care
|January 6, 2025
Summary
Generative artificial intelligence (AI) models show promise in predicting pediatric Emergency Severity Index (ESI) levels. Fine-tuning AI significantly improves its accuracy and reliability for use in pediatric emergency triage.
Area of Science:
- Artificial Intelligence in Healthcare
- Pediatric Emergency Medicine
- Clinical Triage Systems
Background:
- Accurate and timely patient assessment is crucial in pediatric emergency departments.
- The Emergency Severity Index (ESI) is a widely used triage tool.
- Evaluating the potential of generative artificial intelligence (AI) for ESI level prediction in pediatric patients is an emerging area.
Purpose of the Study:
- To assess the accuracy and reliability of multiple generative AI models in predicting pediatric ESI levels.
- To evaluate the impact of medically oriented fine-tuning on AI model performance.
- To compare the performance of various AI models including ChatGPT-3.5, ChatGPT-4.0, T5, Llama-2, Mistral-Large, and Claude-3 Opus.
Main Methods:
- Utilized 70 pediatric clinical vignettes from the ESI Handbook v4 as the gold standard.
- Calculated performance metrics (sensitivity, specificity, F1 score) for each AI model's ESI level predictions.
- Assessed reliability using repeated tests and Fleiss kappa for interrater agreement, comparing models pre- and post-fine-tuning with paired t tests.
Main Results:
- Claude-3 Opus demonstrated the highest performance among untrained models (Sensitivity: 80.6%, Specificity: 91.3%, F1: 73.9%).
- Fine-tuned GPT-4.0 showed significant improvement (Sensitivity: 77.1%, Specificity: 92.5%, F1: 74.6%, P < 0.04).
- High reliability was observed for Claude-3 Opus (κ: 0.85), Mistral-Large (κ: 0.79), and trained GPT-4.0 (κ: 0.67), with training enhancing GPT model reliability (P < 0.001).
Conclusions:
- Generative AI models exhibit notable accuracy in predicting pediatric ESI levels.
- Medically oriented fine-tuning substantially enhances the performance and reliability of these AI models.
- AI holds potential as a valuable adjunct tool for improving pediatric emergency triage.

