Related Experiment Video
Updated: Jan 11, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Exploring the prognostic utility of large language models versus traditional clinical models in heart failure: a
Luxiang Shang1, Yali Chen2, Rui Li3
1Department of Health Management, Shandong Engineering Laboratory of Health Management, Shandong Institute of Health Management, the First Affiliated Hospital of Shandong First Medical University & Shandong Provincial Qianfoshan Hospital, Jinan, China.
Background:
Large language models (LLMs) show promise in clinical decision support; however, their role in risk prediction for heart failure (HF) remains uncertain.
Objective:
This pilot study evaluated the prognostic performance and reproducibility of two general-purpose LLMs, ChatGPT and DeepSeek, using structured clinical data and unstructured discharge summaries, compared with a conventional clinical model.
Methods:
Structured data from the Zigong HF study included 473 hospitalized HF patients with 33 clinical variables predicting a 90-day composite outcome of all-cause death or rehospitalization. Discharge summaries from the MIMIC-IV cohort included 2,091 ICU HF patients predicting 1-year all-cause mortality. Standardized prompts were used to obtain predicted probabilities from each LLM. Model predictions were compared with logistic regression results, and reproducibility was assessed using intraclass correlation coefficients.
Results:
In the Zigong HF study, both LLMs showed limited discrimination (AUC 0.59 for ChatGPT, 0.56 for DeepSeek), performing below the conventional model (AUC 0.63). In the MIMIC-IV cohort, ChatGPT achieved higher discrimination (AUC 0.72) than DeepSeek (AUC 0.67, P < 0.001) and comparable performance to the clinical model (AUC 0.74, P = 0.31). Decision curve analysis showed modest benefit for ChatGPT at low-to-moderate thresholds, while DeepSeek offered minimal benefit. Repeated predictions showed significant variability for both models.
Conclusions:
This pilot study provides preliminary evidence that LLMs have limited predictive value for structured data but show comparable performance in text-based risk prediction. These findings suggest potential for LLMs in processing unstructured clinical information and highlight the need for validation in larger, contemporary cohorts.
Related Concept Videos
Heart Failure IV: Classification and Diagnostic Evaluation
Heart Failure III: Clinical Manifestations
Heart Failure Drugs: Inhibitors of Renin-Angiotensin System
Heart Failure VI: Adjunct Therapies
Heart Failure V: Medical Management
Heart Failure II: Pathophysiology

