Related Experiment Videos
Clinicians vs. Artificial Intelligence in Predicting 28-Day ICU Mortality: A Vignette Study
Suppanut Siriseth1, Veerapong Vattanavanit2
1Division of Internal Medicine, Faculty of Medicine, Prince of Songkla University, Hat Yai, Songkhla, 90110, Thailand.
Context:
The surprise question (SQ) is an intuition-based tool for identifying patients who may benefit from palliative care, but its short-term prognostic performance in critical illness remains uncertain. Large language models may offer standardized prognostic judgments.
Objectives:
To compare healthcare professionals and ChatGPT in predicting 28-day mortality among critically ill patients using the binary 28-day surprise question (SQ-28d).
Methods:
We conducted a multi-rater, vignette-based cross-sectional study in a tertiary medical intensive care unit (ICU) in Thailand. Sixty de-identified clinical vignettes were derived from adult patients with known 28-day outcomes. Consultant intensivists, internal medicine residents, ICU nurses, and ChatGPT each answered the SQ-28d. Primary outcome was correct prognostic classification; secondary outcomes included sensitivity, specificity, predictive values, likelihood ratios, area under the receiver operating characteristic curve (AUROC), and inter-rater agreement.
Results:
. Correct classification differed significantly across groups (P = 0.046): consultant intensivists, 69.4%; residents, 66.4%; nurses, 64.1%; and ChatGPT, 55.0%; however, no pairwise comparison remained significant after Bonferroni adjustment. AUROC was highest for consultant intensivists (0.734, 95% confidence interval [CI]: 0.670-0.798), followed by residents (0.713), nurses (0.657), and ChatGPT (0.657). Consultant intensivists had the highest positive likelihood ratio (2.25) and ChatGPT the lowest negative likelihood ratio (0.14). Inter-rater agreement was substantial between intensivists and residents and moderate between nurses and physicians.
Conclusions:
Healthcare professionals demonstrated higher 28-day mortality prediction accuracy than ChatGPT using the SQ. The SQ-28d may serve as a pragmatic trigger for prognostic awareness and palliative care discussions, but should not be used as a standalone prognostic tool.