Related Experiment Video
Updated: Feb 7, 2026

Expired CO2 Measurement in Intubated or Spontaneously Breathing Patients from the Emergency Department
Published on: January 29, 2011
Development of BERT-based large language models for emergency department triage using real-world conversations
Sukyo Lee1, Sumin Jung2, Jong-Hak Park1
1Department of Emergency Medicine, Korea University Ansan Hospital, Ansan-si, Gyeonggi-do 15355, Republic of Korea.
This study developed BERT-based large language models (LLMs) trained on real emergency department conversations for patient urgency classification. The AI models significantly outperformed ChatGPT in accurately identifying urgent cases.
Area of Science:
- Artificial Intelligence in Medicine
- Clinical Informatics
- Natural Language Processing
Background:
- Accurate emergency department (ED) triage is crucial for effective resource allocation.
- Previous AI models for triage used summarized clinical data, limiting real-world applicability.
- Large language models (LLMs) offer potential for analyzing complex clinical conversations.
Purpose of the Study:
- To develop and evaluate LLMs trained on actual clinical conversations for patient urgency classification.
- To compare the performance of these LLMs against existing AI models and general-purpose chatbots.
- To assess the explainability of the developed LLMs in clinical decision-making.
Main Methods:
- Utilized a curated dataset of anonymized triage conversations from Korean tertiary hospitals.
- Developed two BERT-based LLMs: one tokenizing entire conversations, another using hierarchical structure with speaker-role embeddings.
- Classified patient urgency based on the Korean Triage and Acuity Scale (KTAS) into urgent (KTAS 3) and non-urgent (KTAS 4-5).
- Evaluated models using accuracy, precision, recall, and F1-score, comparing against ChatGPT GPT-4o and ClinicalBERT.
- Assessed model explainability using SHapley Additive exPlanations (SHAP).
Main Results:
- A hierarchical BERT-based LLM achieved 75.94% accuracy, significantly outperforming ChatGPT (56.68%) and ClinicalBERT (69.42%).
- The best model demonstrated a recall of 0.9610 for urgent cases, surpassing ChatGPT's recall of 0.5352.
- SHAP analysis confirmed the model's focus on clinically relevant cues aligned with KTAS criteria.
Conclusions:
- BERT-based LLMs trained on real-world ED conversations provide superior triage accuracy compared to general AI models.
- This approach enhances clinical decision support systems with interpretable and efficient AI.
- The findings highlight the potential of specialized LLMs for improving emergency care workflows.
More Related Videos
09:52Setting Up a Stroke Team Algorithm and Conducting Simulation-based Training in the Emergency Department - A Practical Guide
Published on: January 15, 2017
04:05Author Spotlight: Unveiling Prognostic Indicators in Heart Failure - The Role of Phase Angle and Bioelectrical Impedance Analysis
Published on: June 30, 2023
Related Concept Videos
Language Development
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
Gene Conversion
Gene Conversion
Language
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...
Emerging Adulthood
Components of Language