从大型语言模型中对诊断生成的不确定性估计:下一个词的概率不是测试前的概率
Yanjun Gao1,2, Skatje Myers2, Shan Chen3,4
1Department of Biomedical Informatics, University of Colorado Anschutz Medical Campus, Aurora, CO 80045, United States.
JAMIA open
|January 13, 2025
概括
大型语言模型 (LLM) 显示出诊断概率估计的潜力,但与传统的机器学习分类器相比,目前表现不佳. 需要进一步的研究来提高它们在临床环境中的准确性和可靠性.
科学领域:
- 人工智能在医学中的应用
- 临床决策支持系统 临床决策支持系统
- 机器学习用于医疗保健
背景情况:
- 准确的测试前诊断概率估计对于有效的临床决策至关重要.
- 大型语言模型 (LLM) 正在成为医学数据分析的潜在工具.
- 与已建立的方法对LLM绩效的评估对于其临床采用至关重要.
研究的目的:
- 评估指令调整的LLM在估计前测试诊断概率方面的能力.
- 将LLM的不确定性估计性能与传统的机器学习分类器进行比较.
- 确定在诊断任务的LLM应用中需要改进的领域.
主要方法:
- 评估了两个调整指令的LLM (Mistral-7B-Instruct,Llama3-70B-chat-hf) 的情况.
- 使用电子健康记录 (EHR) 数据对败血症,心律失常和充血性心力衰竭 (CHF) 的二进制结果预测.
- 对比LLM不确定性估计方法 (口头化的信心,令牌逻辑,LLM嵌入+XGB) 与极端梯度增强 (XGB) 分类器.
主要成果:
- 极端梯度增强 (XGB) 分类器在所有基于LLM的方法中表现出优越的性能.
- 在LLM嵌入+XGB方法显示性能最接近基线XGB分类器.
- 语言化信心和令牌逻辑的方法表现明显不佳.
结论:
- 与传统的ML分类器相比,当前的LLM在提供可靠的测试前诊断概率估计方面存在局限性.
- 对于临床环境中的LLM来说,需要改进校准和偏差缓解策略.
- 未来的研究应该专注于混合方法,将LLM与数值推理和校准嵌入整合起来.
相关概念视频
Uncertainty: Confidence Intervals
3.1K
The confidence interval is the range of values around the mean that contains the true mean. It is expressed as a probability percentage. The interpretation of a 95% confidence interval, for instance, is that the statistician is 95% confident that the true mean falls within the interval. The upper and lower limits of this range are known as confidence limits. The confidence limits for the true mean are estimated from the sample's mean, the standard deviation, and the statistical factor...
3.1K
Uncertainty: Overview
509
In analytical chemistry, we often perform repetitive measurements to detect and minimize inaccuracies caused by both determinate and indeterminate errors. Despite the cares we take, the presence of random errors means that repeated measurements almost never have exactly the same magnitude. The collective difference between these measurements - observed values - and the estimated or expected value is called uncertainty. Uncertainty is conventionally written after the estimated or expected value.
509
Propagation of Uncertainty from Systematic Error
462
The atomic mass of an element varies due to the relative ratio of its isotopes. A sample's relative proportion of oxygen isotopes influences its average atomic mass. For instance, if we were to measure the atomic mass of oxygen from a sample, the mass would be a weighted average of the isotopic masses of oxygen in that sample. Since a single sample is not likely to perfectly reflect the true atomic mass of oxygen for all the molecules of oxygen on Earth, the mass we obtain from this...
462
Propagation of Uncertainty from Random Error
636
An experiment often consists of more than a single step. In this case, measurements at each step give rise to uncertainty. Because the measurements occur in successive steps, the uncertainty in one step necessarily contributes to that in the subsequent step. As we perform statistical analysis on these types of experiments, we must learn to account for the propagation of uncertainty from one step to the next. The propagation of uncertainty depends on the type of arithmetic operation performed on...
636
Interpretation of Confidence Intervals
5.6K
A confidence interval is a better estimate of the population than a point estimate, as it uses a range of values from a sample instead of a single value.
Confidence intervals have confidence coefficients that are crucial for their interpretation. The most common confidence coefficients are 0.90, 0.95, and 0.99, which can be written as percentages–90%, 95%, and 99%, respectively.
Suppose a person calculates a confidence interval with a confidence coefficient of 0.95. In that case, they can...
Confidence intervals have confidence coefficients that are crucial for their interpretation. The most common confidence coefficients are 0.90, 0.95, and 0.99, which can be written as percentages–90%, 95%, and 99%, respectively.
Suppose a person calculates a confidence interval with a confidence coefficient of 0.95. In that case, they can...
5.6K
Improving Translational Accuracy
8.6K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
8.6K


