Related Experiment Videos
Generative AI and Clinicians Show Comparable Prognostic Reasoning From Clinical Narratives in Biologic-Treated CRSwNP
Sholem Hack1, Chase Kahn2, Ainhoa Garcia-Lliberos3
1City St. George's University of London, School of Medicine Program delivered by University of Nicosia at the Chaim Sheba Medical Center Ramat Gan Israel.
Objective:
To compare the ability of large language models (LLMs) and otolaryngologists to identify prognostic signals from brief clinical vignettes in CRSwNP.
Methods:
In this blinded study, 68 adults initiating biologic therapy for CRSwNP (≥ 36 months follow-up) were represented by standardized vignettes derived from documentation immediately before biologic initiation. Vignettes summarized symptoms, prior surgery, comorbidities, and medications. Biomarkers, imaging scores, smell testing, and follow-up data were excluded to isolate text-based prognostic reasoning under identical constraints. Five attending otolaryngologists, one rhinology fellow, and two residents independently predicted four 5-year outcomes: subsequent endoscopic sinus surgery, recurrent systemic steroid bursts, biologic switch, and a composite endpoint. Multiple LLMs were queried in identical zero-shot format across three sessions to assess stability. Predictions were compared with verified outcomes using accuracy, sensitivity, specificity, F1 score, Cohen's κ, and Matthews correlation coefficient.
Results:
Five-year outcome prevalences were 38.2% (26 of 68) for subsequent sinus surgery, 26.5% (18 of 68) for steroid bursts, 14.7% (10 of 68) for biologic switch, and 58.8% (40 of 68) for composite failure. Macro-averaged LLM accuracies ranged from 65.0% to 76.1%, compared with 65.7% ± 7.4% for human raters. The best-performing LLM achieved accuracy comparable to the top attending. Both groups showed higher negative than positive predictive values, indicating better discrimination for patients without adverse events. Agreement between LLMs and clinicians was moderate and within the range of inter-clinician variability.
Conclusion:
When restricted to short pretreatment text, LLMs produced prognostic judgments within the range of clinician variability for severe, biologic-treated CRSwNP outcomes. These findings suggest that in information-limited settings, LLMs may provide an auxiliary reasoning signal rather than a standalone decision tool.