Related Experiment Video
Updated: Jun 3, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Uncertainty estimation in diagnosis generation from large language models: next-word probability is not pre-test
Yanjun Gao1,2, Skatje Myers2, Shan Chen3,4
1Department of Biomedical Informatics, University of Colorado Anschutz Medical Campus, Aurora, CO 80045, United States.
Large language models (LLMs) show potential for diagnostic probability estimation but currently underperform compared to traditional machine learning classifiers. Further research is needed to improve their accuracy and reliability in clinical settings.
Area of Science:
- Artificial Intelligence in Medicine
- Clinical Decision Support Systems
- Machine Learning for Healthcare
Background:
- Accurate pre-test diagnostic probability estimation is crucial for effective clinical decision-making.
- Large Language Models (LLMs) are emerging as potential tools for medical data analysis.
- Evaluating LLM performance against established methods is essential for their clinical adoption.
Purpose of the Study:
- To assess the capability of instruction-tuned LLMs in estimating pre-test diagnostic probabilities.
- To compare the uncertainty estimation performance of LLMs with traditional machine learning classifiers.
- To identify areas for improvement in LLM application for diagnostic tasks.
Main Methods:
- Two instruction-tuned LLMs (Mistral-7B-Instruct, Llama3-70B-chat-hf) were evaluated.
- Binary outcome prediction for Sepsis, Arrhythmia, and Congestive Heart Failure (CHF) using electronic health record (EHR) data.
- Comparison of LLM uncertainty estimation methods (Verbalized Confidence, Token Logits, LLM Embedding+XGB) against an eXtreme Gradient Boosting (XGB) classifier.
Main Results:
- The eXtreme Gradient Boosting (XGB) classifier demonstrated superior performance over all LLM-based methods.
- The LLM Embedding+XGB approach showed performance closest to the baseline XGB classifier.
- Verbalized Confidence and Token Logits methods underperformed significantly.
Conclusions:
- Current LLMs have limitations in providing reliable pre-test diagnostic probability estimations compared to traditional ML classifiers.
- Improved calibration and bias mitigation strategies are necessary for LLMs in clinical settings.
- Future research should focus on hybrid approaches integrating LLMs with numerical reasoning and calibrated embeddings.
Related Concept Videos
Uncertainty: Confidence Intervals
Uncertainty: Overview
Propagation of Uncertainty from Systematic Error
Propagation of Uncertainty from Random Error
Interpretation of Confidence Intervals
Confidence intervals have confidence coefficients that are crucial for their interpretation. The most common confidence coefficients are 0.90, 0.95, and 0.99, which can be written as percentages–90%, 95%, and 99%, respectively.
Suppose a person calculates a confidence interval with a confidence coefficient of 0.95. In that case, they can...
Improving Translational Accuracy

