Related Experiment Video
Updated: Aug 30, 2026

Autologous Blood Injection to Model Spontaneous Intracerebral Hemorrhage in Mice
Published on: August 24, 2011
Assessing the feasibility of ChatGPT in predicting outcomes after intracerebral hemorrhage
Noa B Mintz1, Asala Erekat2, Ali Mahta1
1Department of Neurology, Brown University, Alpert Medical School, Providence, RI, United States of America.
Abstract:
Prognostication after intracerebral hemorrhage (ICH) is prone to heterogeneity and bias, which represents a potential opportunity to integrate an artificial intelligence-based approach. We aimed to evaluate the feasibility of a commercially-available large language model in predicting functional outcomes after ICH using routine admission data. We used GPT-4o mini with standardized prompts providing clinical data from the first 24 h of admission to predict 3-month modified Rankin Scale (mRS) in a single-center cohort of 409 ICH patients. We then compared the accuracy of these GPT-predicted scores with patients' actual 3-month outcomes, as well as outcomes predicted by two board-certified neurologists. GPT-predicted scores showed moderate correlation with actual outcomes (r = 0.64 [95% CI: 0.58-0.70]), although they tended to underestimate 3-month mRS (mean difference -0.99 [95% CI: -1.16, -0.83; SD 1.68]). This was similar to the averaged clinician scores (r = 0.68 [95% CI: 0.62-0.73]; mean difference -0.35 [95% CI: -0.49, -0.20; SD 1.47]), although one clinician tended toward underestimation while the other tended toward overestimation. GPT's accuracy in predicting favorable 3-month outcomes (defined as mRS 0-3) was comparable to clinician-averaged scores (AUC 0.86 vs. 0.85), with high sensitivity (88%) and moderate specificity (67%). However, model performance was lower in older patients (r = 0.58 [95% CI: 0.46-0.67]; mean difference -1.44 [95% CI: -1.71, -1.16; SD 1.73]) and those with milder deficits (r = 0.22 [95% CI: 0.06-0.37]; mean difference -1.59 [95% CI: -1.90, -1.28; SD 1.89]). These findings suggest that, although it may underestimate 3-month disability, GPT-4o mini has comparable performance to clinicians and may be a feasible tool to aid prognostication after ICH.

