Related Experiment Video
Updated: Jun 4, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Probabilistic medical predictions of large language models
Bowen Gu1, Rishi J Desai1, Kueiyu Joshua Lin1
1Division of Pharmacoepidemiology and Pharmacoeconomics, Department of Medicine, Brigham and Women's Hospital, Harvard Medical School, Boston, MA, USA.
Large Language Models (LLMs) show promise in clinical predictions but struggle with reliable probability estimates. Implicit probabilities derived from label token likelihood outperform explicit text-generated probabilities in LLM clinical applications.
Area of Science:
- Artificial Intelligence
- Clinical Informatics
- Machine Learning
Background:
- Large Language Models (LLMs) offer flexible clinical prediction capabilities via prompt engineering.
- Reliable prediction probabilities are essential for transparency and informed decision-making in healthcare.
- Current LLMs face challenges in generating trustworthy probability estimates due to numerical reasoning limitations.
Purpose of the Study:
- To compare the reliability of explicit probabilities generated by LLMs through text prompts versus implicit probabilities derived from label token likelihood.
- To evaluate the performance of these probability estimation methods across various LLMs and medical datasets.
- To identify factors influencing the discrepancy between explicit and implicit probability estimations in clinical LLM applications.
Main Methods:
- Evaluated six advanced open-source Large Language Models (LLMs).
- Utilized five diverse medical datasets for performance assessment.
- Compared explicit probability estimates (from text generation) with implicit probability estimates (from correct label token prediction likelihood).
Main Results:
- Implicit probabilities consistently outperformed explicit probabilities across key metrics: discrimination, precision, and recall.
- The performance gap between implicit and explicit probabilities was more significant in smaller LLMs and with imbalanced datasets.
- Explicit probability generation by LLMs demonstrated limitations in clinical prediction reliability.
Conclusions:
- Implicit probability estimation methods show greater reliability for clinical applications compared to explicit methods in current LLMs.
- Findings underscore the need for cautious interpretation of LLM-generated probabilities in healthcare settings.
- Further research is required to develop improved probability estimation techniques for robust clinical deployment of LLMs.
More Related Videos
Related Concept Videos
Language and Cognition
Steps in Outbreak Investigation
Leaky Scanning
Mechanistic Models: Compartment Models in Individual and Population Analysis
Improving Translational Accuracy
Sensitivity, Specificity, and Predicted Value
Sensitivity is the...

