Related Experiment Video
Updated: Jul 13, 2026

09:01
Minimally Invasive Endoscopic Intracerebral Hemorrhage Evacuation
Published on: October 15, 2021
Prompt-Sensitive Decision Behavior of Large Language Models in Intensive Care Unit Mortality Prediction for
Jinn-Rung Kuo1,2,3, Guan-Yu Chen3, Xiao-Han Vivian Yap3
1School of Medicine, College of Medicine, National Sun Yat-Sen University, Kaohsiung, Taiwan.
Journal of Medical Internet Research
|July 9, 2026
Summary
Large language models (LLMs) show moderate predictive performance in clinical risk estimation but exhibit prompt-sensitive behavior. Outcome-trained models offer more reliable quantitative risk assessment for clinical decision support.
Area of Science:
- Clinical prediction modeling
- Artificial intelligence in healthcare
- Medical informatics
Background:
- Large language models (LLMs) are being explored for clinical decision support.
- Reliability of LLM-generated quantitative risk estimates in structured clinical prediction is uncertain.
Purpose of the Study:
- Evaluate LLM predictive performance and decision-making in structured clinical prediction.
- Compare LLM outputs with an outcome-trained machine learning model.
Main Methods:
- Benchmarking study using structured clinical data from intensive care unit patients with intracerebral hemorrhage.
- Compared an extreme gradient boosting model with a general-purpose LLM using various prompting strategies (zero-shot, few-shot, chain-of-thought).
- Evaluated discrimination, classification behavior, and feature importance concordance.
Main Results:
- Outcome-trained model outperformed all LLM approaches in discriminative performance.
- LLMs showed moderate discrimination but significant variability in classification behavior across prompts.
- LLM optimal thresholds were substantially higher (0.74-0.88) than the machine learning model (0.1555).
- Modest concordance between LLM feature prioritization and model-derived feature importance.
Conclusions:
- Inference-only LLM outputs in structured clinical prediction are prompt-sensitive and require cautious interpretation for risk estimation.
- Integrating outcome-trained models with LLM-assisted reasoning may enhance future clinical decision support systems.

