使用具有指令调整的大型语言模型,用于可靠的临床脆弱性评分
Xiang Lee Jamie Kee1, Gerald Gui Ren Sng2,3, Daniel Yan Zheng Lim3,4
1Department of Geriatric Medicine, Singapore General Hospital, Singapore, Singapore.
Journal of the American Geriatrics Society
|August 6, 2024
概括
人工智能,特别是大型语言模型 (LLM),显示出可靠的临床虚弱度量表 (CFS) 评分的希望. 调整到指令的提示提升了LLM的一致性,有可能改善医疗保健中的脆弱性评估.
科学领域:
- 老年学是指老年学的学科.
- 人工智能在医学中的应用
背景情况:
- 脆弱性是健康结果的关键预测因素,其特征是生理衰退和增加的脆弱性.
- 临床脆弱度量 (CFS) 是广泛使用的,但容易受到评级者偏见的影响.
- 大型语言模型 (LLM) 为客观和高效的脆弱性评估提供了一个潜在的解决方案.
研究的目的:
- 评估基于LLM的临床脆弱性量表 (CFS) 评分的可靠性和一致性.
- 为了比较基本与指令调整提示的性能,用于LLM脆弱性评估.
- 为了评估与人类评分器基准对比的LLM评分准确性.
主要方法:
- 利用七个标准化的患者场景进行LLM评估.
- 测试了两个提示方法:基本和指令调整 (包括CFS定义和温度控制).
- 雇员曼 - 惠特尼U测试和弗莱斯卡帕测试用于统计分析和评审者之间的可靠性,比较LLM输出与历史的人类分数.
主要成果:
- 士学位的中位数与人类评分器密切一致 (在1分之内).
- 与基本提示相比,与指令调整的提示显著提高了在五种场景中的分数分配一致性.
- 调整为指令的LLM表现出高的评分间可靠性 (Fleiss' Kappa = 0.887) 和一致的得分.
结论:
- LLM显示出可靠和一致的临床脆弱性评分的巨大潜力.
- 快速工程,特别是指令调整,是优化医疗保健LLM绩效的有效方法.
- LLM可能会提供更客观的脆弱性得分,特别是当详细的患者数据有限时,这需要在临床应用中进行进一步的研究.
相关概念视频
Stereotype Content Model
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence categorization, a person will feel...
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
This lesson introduces two critical methods in pharmacokinetics, the Wagner-Nelson and Loo-Riegelman methods, used for estimating the absorption rate constant (ka) for drugs administered via non-intravenous routes. The Wagner-Nelson method relates ka to the plasma concentration derived from the slope of a semilog percent unabsorbed time plot. However, it is limited to drugs with one-compartment kinetics and can be impacted by factors like gastrointestinal motility or enzymatic degradation.
On...
On...
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
Mechanistic models play a crucial role in algorithms for numerical problem-solving, particularly in nonlinear mixed effects modeling (NMEM). These models aim to minimize specific objective functions by evaluating various parameter estimates, leading to the development of systematic algorithms. In some cases, linearization techniques approximate the model using linear equations.
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
Self-Evaluation Maintenance Model
The Self-Evaluation Maintenance (SEM) model offers a psychological framework to understand how individuals’ self-esteem is influenced by the achievements of others, particularly those with whom they share close personal bonds. The SEM model operates when personal rather than social identity guides individuals. Central to this model is the notion that individuals have an inherent desire to preserve a favorable self-image, which is continuously shaped by interpersonal comparisons and...


