评估和提高大型语言模型在特定领域医学中的表现:使用DocOA进行开发和可用性研究
Xi Chen1,2, Li Wang1,2, MingKe You1,2
1Sports Medicine Center, West China Hospital, Sichuan University, Chengdu, China.
Journal of medical Internet research
|June 4, 2024
概括
像DocOA这样的专门的大型语言模型 (LLM) 在骨关节炎管理方面优于一般的LLM. 这项研究为评估医学LLM引入了新的基准,表明量身定制的方法可以增强临床应用.
科学领域:
- 人工智能在医学中的应用
- 自然语言处理自然语言处理.
- 临床决策支持 临床决策支持
背景情况:
- 大型语言模型 (LLM) 在专业医疗领域,如骨关节炎 (OA) 管理中的有效性在很大程度上是未知的.
- 一般的LLM可能缺乏复杂疾病管理所需的精度.
- 需要为医学LLM提供强大的评估框架.
研究的目的:
- 评估和增强在特定医疗领域的LLMs的临床能力和可解释性.
- 使用骨关节炎 (OA) 管理作为LLM申请的案例研究.
- 开发和验证医学LLM领域特定的基准.
主要方法:
- 开发了一个特定领域的基准框架,用于评估从知识到临床应用的LLM.
- 创建了DocOA,一个专门的LLM用于OA管理,使用检索增强生成和指令提示.
- 使用客观和人类评估在现实世界的临床场景中比较了GPT-3.5,GPT-4和DocOA.
主要成果:
- 一般LLM (GPT-3.5,GPT-4) 在OA管理中表现出有限的有效性,特别是在个性化治疗建议方面.
- 专业的LLM,DocOA,表现显著改善了业绩.
- DocOA的检索增强生成使得临床证据的识别成为可能,提高了答案的解释性.
结论:
- 引入了一个新的基准,以全面评估特定领域的LLM能力.
- 突出了临床环境中普遍的LLM的局限性.
- 展示了定制的LLM方法的潜力,以开发有效的特定领域的医疗AI工具.
相关概念视频
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
Improving Translational Accuracy
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...


