Related Experiment Video
Updated: May 5, 2026

Long-term Continuous EEG Monitoring in Small Rodent Models of Human Disease Using the Epoch Wireless Transmitter System
Published on: July 21, 2015
Can the large language model ChatGPT-4omni predict outcomes in adult patients with status epilepticus?
Simon A Amacher1,2, Sira M Baumann1, Sebastian Berger1
1Clinic for Intensive Care Medicine, Department of Acute Care, University Hospital Basel, Basel, Switzerland.
Objective:
Large language models (LLMs) have recently gained attention for clinical decision-making and diagnosis. This study evaluates the performance of the recently updated LLM Chat Generative Pre-Trained Transformer-4omni (ChatGPT-4o) in predicting clinical outcomes in patients with status epilepticus and compares its prognostic performance to the Status Epilepticus Severity Score (STESS).
Methods:
This retrospective single-center cohort study was performed at the University Hospital Basel (tertiary academic medical center) from January 2005 to December 2022. It included consecutive adult patients (≥18 years of age) with a diagnosis of status epilepticus. The primary outcome was survival at hospital discharge, and the secondary outcome was return to premorbid neurological function at hospital discharge. The performance characteristics of ChatGPT4-o (sensitivity, specificity, Youden Index) were evaluated and compared to those of the STESS.
Results:
Of 760 patients, 689 patients (90.7%) survived to discharge, and 317 survivors (41.7%) regained their premorbid neurological function at discharge. ChatGPT-4o predicted survival in 567 of 760 patients (74.6%), of which 45 died. ChatGPT-4o predicted death in 193 of 760 patients (25.4%), of which 167 survived, resulting in a sensitivity of 75.8% and a specificity of 36.6% (Youden Index 0.12, 95% confidence interval [CI] 0-.28) for predicting survival. ChatGPT-4o predicted return to premorbid neurologic function in 249 of 760 patients (32.8%), of which 112 did not return to their premorbid neurological function. ChatGPT-4o predicted no return to premorbid function in 511 of 760 patients (67.2%), of which 180 returned to their premorbid function, resulting in a sensitivity of 43.2% and a specificity of 74.7% (Youden Index .12, 95% CI .08-.28) for predicting return to premorbid neurological function. There was no difference in the prognostic performance of ChatGPT-4o and the STESS. A second round of prompting did not increase the predictive performance of ChatGPT-4o.
Significance:
ChatGPT-4o unreliably predicts outcomes in patients with status epilepticus. Clinicians should refrain from using ChatGPT-4o for prognostication in these patients.
More Related Videos
Related Concept Videos
Epilepsy and Seizures: Overview
Various factors can trigger epilepsy, including genetic factors, brain damage, metabolic causes, and unknown etiology. Diagnosis of epilepsy involves electroencephalography (EEG), which...
Seizures: Classification
Seizures are typically classified into two main categories: focal and generalized seizures.
Focal Seizures
Focal seizures originate from specific regions of the brain. These seizures are further sub-classified into two types:
Seizures l: Introduction
Epilepsy ll: Types

