在生物医学NLP中的基准检索检索增强的大型语言模型:应用,稳定性和自我意识
Mingchen Li1, Zaifu Zhan2, Han Yang3
1Division of Computational Health Sciences, Department of Surgery University of Minnesota, Minneapolis, MN, USA.
Science advances
|November 21, 2025
概括
检索增强的大型语言模型 (LLM) 在生物医学自然语言处理 (NLP) 任务中表现有前途. 然而,它们需要进一步开发,以提高复杂场景中的稳定性和自我意识.
科学领域:
- 生物医学自然语言处理 (NLP)
- 人工智能 (AI) 是一种人工智能.
- 机器学习 (ML) 是指机器学习.
背景情况:
- 大型语言模型 (LLM) 可以产生幻觉.
- 检索增强的LLM (RALs) 通过检索外部知识来减轻幻觉.
- 在生物医学NLP任务中RALs的有效性尚未得到充分证实.
研究的目的:
- 引入一个全面的基准来评估生物医学NLP中的RAL.
- 评估RAL在各种任务和强度测试台上的表现.
- 提出方法来提高RAL的稳定性和负面意识.
主要方法:
- 开发了生物医学检索增强代基准 (BARGE).
- 在5个生物医学NLP任务和11个数据集上评估了RAL.
- 使用了四个测试台:无标记,反事实,多样化的强度和自我意识.
- 提出了检测和纠正策略和对比学习以改进.
主要成果:
- 一般来说,RAL在生物医学NLP中表现优于标准LLM.
- RAL在稳定性和自我意识方面表现出局限性,特别是在反事实和多样化的场景中.
- 拟议的方法显著提高了对未标记和反事实数据的稳定性.
- 改进了模型检测和避免错误预测的能力.
结论:
- 目前的RAL显示出潜力,但需要对生物医学应用进行改进.
- 坚固性和自我意识仍然是RAL在医疗保健中的关键挑战.
- 需要进一步的研究,以确保RAL在高风险的生物医学环境中的可靠性和准确性.
更多相关视频
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
1.3K
09:20Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications
Published on: February 23, 2019
9.2K
相关概念视频
Improving Translational Accuracy
14.0K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.0K
Improving Translational Accuracy
3.5K
3.5K
