Related Experiment Video
Updated: Jan 9, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
994
Uncertainty Quantification for Clinical Outcome Predictions with (Large) Language Models
Zizhang Chen1, Peizhao Li2, Xiaomeng Dong2
1Brandeis University.
Summary
This study enhances AI reliability in healthcare by quantifying uncertainty in language models (LMs) for electronic health records (EHRs). Methods like ensembling and multi-tasking reduce prediction uncertainty, improving AI transparency and patient safety.
Area of Science:
- Artificial Intelligence in Medicine
- Clinical Informatics
- Machine Learning for Healthcare
Background:
- Language models (LMs) show promise for clinical prediction using electronic health records (EHRs).
- High-stakes healthcare applications demand reliable AI predictions, necessitating robust uncertainty quantification.
- Current AI models often lack transparency, posing risks to patient safety and ethical standards.
Purpose of the Study:
- To develop and validate a framework for uncertainty quantification of LMs in EHR tasks.
- To address uncertainty in both white-box (accessible parameters) and black-box (proprietary LMs like GPT-4) settings.
- To enhance the reliability and transparency of AI-driven clinical predictions.
Main Methods:
- Quantified uncertainty in white-box LMs using multi-tasking and ensemble techniques.
- Extended uncertainty quantification to black-box models, including proprietary LMs.
- Validated the framework on longitudinal clinical data from over 6,000 patients across ten prediction tasks.
Main Results:
- Proposed multi-tasking and ensemble methods effectively reduced model uncertainty in EHR tasks.
- Ensembling and multi-task prediction prompts demonstrated uncertainty reduction across various clinical prediction scenarios.
- The framework successfully increased model transparency in both white-box and black-box settings.
Conclusions:
- Uncertainty quantification using ensembling and multi-tasking improves the reliability of LMs for EHRs.
- The developed framework enhances AI transparency and trustworthiness in clinical decision support.
- This work advances the safe and ethical integration of AI in healthcare delivery.
Related Concept Videos
Uncertainty: Confidence Intervals
10.1K
The confidence interval is the range of values around the mean that contains the true mean. It is expressed as a probability percentage. The interpretation of a 95% confidence interval, for instance, is that the statistician is 95% confident that the true mean falls within the interval. The upper and lower limits of this range are known as confidence limits. The confidence limits for the true mean are estimated from the sample's mean, the standard deviation, and the statistical factor...
10.1K
Uncertainty: Overview
1.5K
In analytical chemistry, we often perform repetitive measurements to detect and minimize inaccuracies caused by both determinate and indeterminate errors. Despite the cares we take, the presence of random errors means that repeated measurements almost never have exactly the same magnitude. The collective difference between these measurements - observed values - and the estimated or expected value is called uncertainty. Uncertainty is conventionally written after the estimated or expected value.
1.5K
Prediction Intervals
3.1K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
3.1K
Propagation of Uncertainty from Random Error
1.6K
An experiment often consists of more than a single step. In this case, measurements at each step give rise to uncertainty. Because the measurements occur in successive steps, the uncertainty in one step necessarily contributes to that in the subsequent step. As we perform statistical analysis on these types of experiments, we must learn to account for the propagation of uncertainty from one step to the next. The propagation of uncertainty depends on the type of arithmetic operation performed on...
1.6K
Improving Translational Accuracy
14.0K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.0K
Improving Translational Accuracy
3.5K
3.5K

