Related Experiment Video
Updated: Jun 20, 2026

A Bedside, Single Burr Hole Approach to Multimodality Monitoring in Severe Brain Injury
Published on: March 26, 2019
Multimodal GPT-5 for Predicting Poor Functional Outcomes After Intracerebral Hemorrhage in the Emergency Department:
Koutarou Matsumoto1,2, Kazuaki Ishihara3, Ryota Tamba4
1Department of Health Care Administration and Management, Graduate School of Medical Sciences, Kyushu University, 3-1-1 Maidashi, Higashi-ku, Fukuoka-shi, Fukuoka, 812-8582, Japan, 81 0926426960.
Large language models like GPT demonstrated comparable discrimination to machine learning for predicting poor outcomes in intracerebral hemorrhage (ICH). However, GPT models showed limitations in calibration and overall accuracy, suggesting a complementary role in clinical decision support.
Area of Science:
- Medical Artificial Intelligence
- Clinical Informatics
- Neurology
Background:
- Rapid prognostic assessment of intracerebral hemorrhage (ICH) is crucial in emergency settings.
- Large language models (LLMs) show potential as clinical decision-support tools.
Purpose of the Study:
- Evaluate the predictive performance of GPT (OpenAI)-based models for poor functional outcomes after ICH.
- Assess the clinical utility of GPT models using real-world multimodal data.
Main Methods:
- Compared GPT-4.1 and GPT-5 with a conventional machine learning (ML) model using clinical and CT imaging data.
- Evaluated models on discrimination (AUROC), overall performance (Brier score, R²), calibration, reproducibility (ICC), and clinical utility (decision curve analysis).
Main Results:
- GPT models showed discrimination comparable to the ML model but inferior overall performance and calibration.
- Model-informed prompting improved GPT discrimination and reproducibility but did not overcome calibration limitations.
- Decision curve analysis indicated limited clinical utility for GPT models compared to the ML model.
Conclusions:
- GPT models achieved comparable discrimination to ML for ICH prognosis but had calibration and accuracy limitations.
- GPT-based models may serve as complementary tools, translating predictive outputs into natural language for enhanced decision-making.

