大型语言模型的概率医学预测
Bowen Gu1, Rishi J Desai1, Kueiyu Joshua Lin1
1Division of Pharmacoepidemiology and Pharmacoeconomics, Department of Medicine, Brigham and Women's Hospital, Harvard Medical School, Boston, MA, USA.
NPJ digital medicine
|December 20, 2024
概括
大型语言模型 (LLM) 在临床预测方面表现有前途,但在可靠的概率估计方面存在困难. 在LLM临床应用中,从标签令牌概率中获得的隐性概率优于明确文本生成的概率.
科学领域:
- 人工智能的人工智能
- 临床信息学 临床信息学
- 机器学习 机器学习
背景情况:
- 大型语言模型 (LLM) 通过快速工程提供灵活的临床预测能力.
- 可靠的预测概率对于医疗保健的透明度和知情决策至关重要.
- 由于数值推理的局限性,目前的LLM在生成可靠的概率估计方面面临着挑战.
研究的目的:
- 为了比较LLM通过文本提示生成的明确概率的可靠性与来自标签令牌概率的隐性概率的可靠性.
- 在各种LLM和医疗数据集中评估这些概率估计方法的性能.
- 确定影响临床LLM应用中的明确和隐含概率估计之间的差异的因素.
主要方法:
- 评估了六个先进的开源大型语言模型 (LLM).
- 利用五个不同的医疗数据集进行绩效评估.
- 将明确的概率估计 (来自文本生成) 与隐性概率估计 (来自正确的标签令牌预测概率) 进行比较.
主要成果:
- 隐式概率在关键指标 (歧视,精度和回忆) 中始终优于显式概率.
- 隐式和显式概率之间的性能差距在较小的LLM和不平衡数据集中更为显著.
- 通过LLM显式概率生成表明临床预测可靠性的局限性.
结论:
- 隐式概率估计方法在临床应用中显示出更高的可靠性,而不是当前LLMs中的显式方法.
- 这些发现强调了在医疗保健环境中对LLM产生的概率进行谨慎解释的需要.
- 需要进一步的研究来开发改进的概率估计技术,以便对LLMs进行强大的临床部署.
相关概念视频
Language and Cognition
322
Language serves as a bridge between ideas and communication, influencing how individuals perceive and interact with the world. Psychologists have long debated whether language shapes thought or vice versa. This discussion gained grip with Edward Sapir and Benjamin Lee Whorf in the 1940s, who proposed that language determines thought, a concept known as linguistic determinism. They suggested that the vocabulary and structure of a language influence how its speakers think and perceive reality.
322
Steps in Outbreak Investigation
105
In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
105
Leaky Scanning
5.1K
During most eukaryotic translation processes, the small 40S ribosome subunit scans an mRNA from its 5' end until it encounters the first start AUG codon. The large 60S ribosomal subunit then joins the smaller one to initiate protein synthesis. The location of the translation initiation is largely determined by the nucleotides near the start codon as there may be multiple translation initiation sites present on the mRNA. Marilyn Kozak discovered that the sequence RCCAUGG (where R...
5.1K
Mechanistic Models: Compartment Models in Individual and Population Analysis
27
Mechanistic models are utilized in individual analysis using single-source data, but imperfections arise due to data collection errors, preventing perfect prediction of observed data. The mathematical equation involves known values (Xi), observed concentrations (Ci), measurement errors (εi), model parameters (ϕj), and the related function (ƒi) for i number of values. Different least-squares metrics quantify differences between predicted and observed values. The ordinary least...
27
Improving Translational Accuracy
2.5K
2.5K
Sensitivity, Specificity, and Predicted Value
180
In healthcare diagnostics, laboratory tests play a crucial role in identifying and diagnosing a wide range of medical conditions. However, interpreting test results is not always straightforward. An abnormal test result does not always confirm the presence of a disease, just as a normal result does not guarantee its absence. To assess the reliability of these diagnostic tools, healthcare practitioners rely on two key statistical indicators: sensitivity and specificity.
Sensitivity is the...
Sensitivity is the...
180


