概括
大型语言模型显示出改善的临床推理和多式联络式问答能力,但在具体任务上却很难. 当前的基准标准可能会高估能力,需要对安全的临床部署进行仔细的人在循环评估.
科学领域:
- 人工智能在医学中的应用
- 生物医学信息学 生物医学信息学
- 自然语言处理自然语言处理.
背景情况:
- 大型语言模型 (LLM) 接近临床使用,但它们的可靠性和基准有效性需要彻底检查.
- 评估LLM在现实世界生物医学任务中的表现对于安全有效的部署至关重要.
研究的目的:
- 在生物医学文本挖掘和问答任务上对边境LLM进行全面,可重复的审计.
- 评估各种环境中的LLM绩效,包括推理密集型,提取型和多式联络型任务.
- 评估当前基准的适用性,并确定模型能力的潜在错误估计.
主要方法:
- 一个统一的,以人为中心的审计领先的通用LLMs.
- 使用了代表性的生物医学文本挖掘任务和九个生物医学问答基准.
- 包括盲目的专家判断,以评估临床合理性和推理一致性.
主要成果:
- 在临床推理和多式联络生物医学问答方面,LLMs表现出持续的改进.
- 在格式受限的任务中仍然存在挑战,例如跨度层次提取和证据密集的总结.
- 专家审查显示,基准标注注释错误可能误估LLM的能力.
- 成本正常化分析显示,对于最近的模型来说,以更低的成本提高了准确性.
结论:
- 一般用途的LLM正在接近临床应用的部署相关可靠性.
- 具体任务的局限性需要混合架构和循环中的人类系统.
- 安全有效的临床整合需要持续的专家监督和评估.
相关概念视频
Improving Translational Accuracy
15.3K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
15.3K
Improving Translational Accuracy
3.7K
3.7K
Reasoning
480
Reasoning is the action of thinking about something in a logical, sensible way. It is integral to problem-solving, decision-making, and critical thinking. Reasoning can be inductive or deductive. Reasoning involves transforming information into conclusions, which is essential for problem-solving, decision-making, and critical thinking.
Inductive reasoning involves deriving generalizations from specific observations. This type of reasoning helps form beliefs about the world. For example,...
Inductive reasoning involves deriving generalizations from specific observations. This type of reasoning helps form beliefs about the world. For example,...
480

