Related Experiment Video
Updated: Jan 11, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Piloting Temperature-Driven Variability in Emergency Diagnostic Accuracy Using a Leading Large Language Model
Philip C Jarrett1, Jared Hill1, Marshall Howell1
1Emergency Medicine, University of Texas Southwestern Medical Center, Dallas, USA.
Abstract:
Background Large language models (LLMs) use a parameter known as temperature to control the stochasticity of output sampling during text outputs, which may have implications for clinical diagnostic tasks. In this study, we aimed to determine the impact of the temperature parameter on GPT-4o's diagnostic accuracy when evaluating emergency medicine cases and assess the effect on diagnostic divergence across iterations. Methodology We conducted a simulation-based diagnostic accuracy study using four challenging emergency medicine cases adapted from the Foundations of Emergency Medicine curriculum. Each case was submitted to GPT-4o 250 times at five temperature settings (0.0, 0.25, 0.50, 0.75, 1.0), both with and without physical examination findings, yielding 10,000 total outputs. Each output contained exactly three differential diagnoses with one leading diagnosis to limit the inflation of diagnostic accuracy by larger, unprioritized lists of remotely possible diagnoses. Diagnostic accuracy was assessed by comparing outputs against predetermined gold-standard diagnoses. Mixed-effects models evaluated the relationship between temperature and diagnostic accuracy, while a sensitivity analysis excluded physical examination data. Diagnostic divergence, defined as the number of unique diagnoses generated across iterations, was explored within cases as a representation of internal consistency. Results At temperature 0.0, GPT-4o achieved 100% leading diagnosis accuracy across all cases with physical examination data. As the temperature increased, the accuracy declined systematically to 89.4% at the temperature setting of 1.0. Mixed-effects models demonstrated that temperature was inversely associated with correct leading diagnosis (β = -4.16, odds ratio (OR) = 0.02, 95% confidence interval (CI) = 0.01-0.03, p < 0.001) and with inclusion of the gold-standard diagnosis anywhere in the differential (β = -3.75, OR = 0.02, 95% CI = 0.01-0.05, p < 0.001). Diagnostic divergence increased from an average of 4.5 unique diagnoses at temperature 0.0 to 26.25 at temperature 1.0 (483% increase). Case sensitivity varied significantly, with ascending cholangitis showing the greatest temperature sensitivity (accuracy dropping from 100% to 70.4%), while carbon monoxide poisoning maintained 100% accuracy across all settings. Sensitivity analysis evaluating the impact of physical examination on diagnostic accuracy revealed case-specific effects. While the diagnostic accuracies of ascending cholangitis and myxedema coma were heavily affected by the exclusion of physical examination data, the carbon monoxide and cryptococcal meningitis cases were minimally changed, if at all. Conclusions Increasing the GPT-4o temperature parameter systematically introduced diagnostic inaccuracy across four emergency medicine vignettes. Lower temperature settings led to improved diagnostic accuracy and consistency across case iterations, which may make them preferable for clinical applications requiring high reliability. Transparent reporting of temperature settings is essential for reproducible clinical artificial intelligence research.
More Related Videos
09:09Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy
Detection of Gross Error: The Q Test
Variability: Analysis
The range is a simple measure of variability, indicating the difference between the highest and...
Random Error
Survival Tree
Building a Survival Tree
Constructing a...