Related Experiment Video
Updated: Sep 9, 2025

Setting Up a Stroke Team Algorithm and Conducting Simulation-based Training in the Emergency Department - A Practical Guide
Published on: January 15, 2017
Performance of ChatGPT, Gemini and DeepSeek for non-critical triage support using real-world conversations in
Sukyo Lee1, Sumin Jung2, Jong-Hak Park1
1Department of Emergency Medicine, Korea University Ansan Hospital, Ansan-si, 15355, Republic of Korea.
Background:
Timely and accurate triage is crucial for the emergency department (ED) care. Recently, there has been growing interest in applying large language models (LLMs) to support triage decision-making. However, most existing studies have evaluated these models using simulated scenarios rather than real-world clinical cases. Therefore, we evaluated the performance of multiple commercial LLMs for non-critical triage support in ED using real-world clinical conversations.
Methods:
We retrospectively analyzed real-world triage conversations prospectively collected from three tertiary hospitals in South Korea. Multiple commercial LLMs-including OpenAI GPT-4o, GPT-4.1, O3, Google Gemini 2.0 flash, Gemini 2.5 flash, Gemini 2.5 pro, DeepSeek V3, and DeepSeek R1-were evaluated for the accuracy in triaging patient urgency based solely on unsummarized dialogue. The Korean Triage and Acuity Scale (KTAS) assigned by triage nurses was used as the gold standard for evaluating the LLM classifications. Model performance was assessed under both a zero-shot prompting condition and a few-shot prompting condition that included representative examples.
Results:
A total of 1,057 triage cases were included in the analysis. Among the models, Gemini 2.5 flash achieved the highest accuracy (73.8%), specificity (88.9%), and PPV (94.0%). Gemini 2.5 pro demonstrated the highest sensitivity (90.9%) and F1-score (82.4%), though with lower specificity (23.3%). GPT-4.1 also showed balanced high accuracy (70.6%) and sensitivity (81.3%) with practical response times (1.79s). Performance varied widely between models and even between different versions from the same vendor. With few-shot prompting, most models showed further improvements in accuracy and F1-score.
Conclusions:
LLMs can accurately triage ED patient urgency using real-world clinical conversations. Several models demonstrated both high sensitivity and acceptable response times, supporting the feasibility of LLM in non-critical triage support tools in diverse clinical environments. These findings apply to non-critical patients (KTAS 3-5), and further research should address integration with objective clinical data and real-time workflow.
Related Concept Videos
Pulmonary Embolism II: Diagnostic Studies and Interprofessional Care
Types of Reports III: Telephone and Verbal Reports
Here's an overview of each type:
Telephone Orders
Aneurysm III: Interprofessional Care
Acute Kidney Injury V: Interprofessional Care
Interdisciplinary Care: The Health Care Team-II
Physical Therapist
A physical therapist (PT) aims to restore function or prevent additional impairment in a patient following an injury or disease. Massage, heat, cold, water, sonar waves, exercises, and electrical stimulation are some treatments used by PTs to treat...
Techniques of Therapeutic Communication II: Focusing, Paraphrasing, and Summarizing
This therapeutic technique can also be used when a patient brings up pertinent information during a health-related conversation. The...

