Related Experiment Video
Updated: Feb 26, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Simulated evaluation of large language model stepwise diagnostic reasoning with real-world chest pain encounters and
Conrad W Safranek1,2, Vimig Socrates1,3, Donald Wright1,2
1Department of Biomedical Informatics and Data Science, School of Medicine, Yale University, New Haven, CT, USA.
Background:
Real-world evaluation of large language models (LLMs) as clinical diagnostic aids is limited by the reliance on static vignettes and retrospective data, which inadequately reflect the dynamic, iterative nature of clinical decision-making and may overestimate LLMs' performance. Here, we benchmark GPT-4o in a stepwise simulated diagnostic setting with real-world clinical data, comparing its diagnostic accuracy and information-seeking strategy with Bayesian-network-derived optimal policies and observed physician practice.
Methods:
We assessed GPT-4o across 500 emergency department (ED) chest-pain encounters, drawn from a cohort of 202,632 cases spanning three EDs. A Bayesian network (BN) trained on the structured cohort data imputed clinical data not collected in the original encounter to create a more robust simulation environment. The BN furthermore enabled derivation of mutual-information-optimal query pathways. GPT-4o sequentially requested information from 136 structured clinical variables under three prompting regimes that varied in disease-prevalence cues and diagnostic category constraints. Diagnostic decisions encompassed one of seven predefined emergent conditions or Other Diagnosis. We measured diagnostic accuracy under each prompting strategy, as well as calculated rank-based overlap with the BN optimal pathway to benchmark the LLM's information-seeking behavior.
Results:
Across the full chest-pain cohort, life-threatening etiologies accounted for only 2.14% of encounters (from 1.04% acute coronary syndrome to 0.01% esophageal rupture). With baseline prompting, GPT-4o systematically over-predicted rare conditions (sensitivity 79.3%; specificity 45.2%); adding prevalence cues or removing diagnostic category constraints respectively increased specificity (83.0% and 94.7%) while reducing false alarms by 107 and 140 per 500, but at the cost of poor sensitivity (30.4% and 8.8%). Rank-biased overlap between GPT-4o's information-seeking sequence and the Bayesian-network mutual-information optimum was low across diagnoses (range 0.060-0.097), and the model diverged from clinician behavior by requesting fewer vitals ([Formula: see text]-fold) and labs ([Formula: see text]-fold), while requesting 30%+ more imaging data.
Conclusions:
In this simulated assessment, GPT-4o demonstrated diagnostic biases toward rare conditions and differed substantially from normative probabilistic models and physician practice patterns. These discrepancies could lead to unnecessary over-triage and resource utilization. Integrating LLMs within more rigorous probabilistic frameworks and calibrating them to realistic disease prevalences may be essential for effectively harnessing their potential as clinical decision-support tools.
Related Concept Videos
Formulating and Validating Nursing Diagnosis II
Risk nursing diagnoses represent clinical judgments of an individual, family, or community more vulnerable to developing the health problem than others...
Formulating and Validating Nursing Diagnosis I
There are thirteen domains...
Documentation of Nursing Diagnosis
In some settings, data-driven computerized decision support systems are in place, allowing for more accurate nursing diagnoses. The database within one of these systems includes diagnostic labels defining characteristics, activities, and indicators for nursing. A nurse enters...
Angina III: Clinical Manifestations and Assessment
Classification of Illness
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
