Related Experiment Video
Updated: Feb 7, 2026

Expired CO2 Measurement in Intubated or Spontaneously Breathing Patients from the Emergency Department
Published on: January 29, 2011
Development of BERT-based large language models for emergency department triage using real-world conversations
Sukyo Lee1, Sumin Jung2, Jong-Hak Park1
1Department of Emergency Medicine, Korea University Ansan Hospital, Ansan-si, Gyeonggi-do 15355, Republic of Korea.
Objectives:
Accurate triage in emergency departments (ED) is critical for appropriate resource allocation. While artificial intelligence (AI) has been explored for triage, prior models relied on summarized clinical scenarios. We aimed to develop and evaluate large language models (LLMs) trained on real-world clinical conversations to classify patient urgency.
Materials And Methods:
We used a nationally curated dataset of anonymized triage-level conversations from 3 tertiary Korean hospitals. Two BERT-based models were developed to classify urgency per the Korean Triage and Acuity Scale (KTAS) into urgent (KTAS 3) or non-urgent (KTAS 4-5). One model tokenized the entire conversation, while the other applied a hierarchical structure with sentence-level tokenization and speaker-role embeddings. Performance metrics included accuracy, precision, recall, and F1-score. We compared our models against ChatGPT GPT-4o and ClinicalBERT, and assessed explainability using SHapley Additive exPlanations (SHAP).
Results:
A total of 5244 clinical conversations, 1057 triage-level dialogues were used, with 950 for training and 107 for testing. Our model with hierarchical structure achieved accuracies of 75.94%, significantly outperforming ChatGPT (56.68%) or fine-tuned ClinicalBERT (69.42%). For urgent cases, the best model achieved a recall of 0.9610, outperforming ChatGPT (0.5352). SHapley Additive exPlanations analysis confirmed that our model focused on clinically relevant cues aligned with KTAS criteria.
Conclusion:
BERT-based LLMs trained on real-world ED conversations significantly outperform general-purpose models like ChatGPT in triage accuracy. This approach demonstrates the potential for enhancing clinical decision support with interpretable and efficient AI.
More Related Videos
09:52Setting Up a Stroke Team Algorithm and Conducting Simulation-based Training in the Emergency Department - A Practical Guide
Published on: January 15, 2017
04:05Author Spotlight: Unveiling Prognostic Indicators in Heart Failure - The Role of Phase Angle and Bioelectrical Impedance Analysis
Published on: June 30, 2023
Related Concept Videos
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Gene Conversion
Gene Conversion
Language
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...
Emerging Adulthood
Components of Language