评估生物医学微调对临床任务的大型语言模型的有效性
Felix J Dorfner1,2, Amin Dada3, Felix Busch4
1Charité-Universitätsmedizin Berlin, Corporate Member of Freie Universität Berlin and Humboldt-Universität zu Berlin, Berlin 10117, Germany.
概括
生物医学大语言模型 (LLM) 在临床任务上往往表现不佳,相比一般用途的LLM. 对医疗数据进行微调的LLM可能不会像临床应用中检索增强那样有效.
科学领域:
- 人工智能在医学中的应用
- 自然语言处理自然语言处理.
- 临床信息学 临床信息学
背景情况:
- 大型语言模型 (LLM) 对生物医学应用有希望.
- 特定领域的微调是提高LLM绩效的常见策略.
- 对LLM生物医学微调的实际好处仍然不确定.
研究的目的:
- 批判性地评估生物医学上微调的LLMs的表现.
- 将它们与各种临床任务中的通用LLM进行比较.
- 评估微调模型对未见数据的概括能力.
主要方法:
- 评估生物医学上微调的LLM与通用LLM相比.
- 使用了NEJM和JAMA的临床案例挑战.
- 评估信息提取,总结和临床编码的性能,使用微调数据集之外的各种基准.
主要成果:
- 生物医学LLM通常表现低于通用模型,特别是在非医学知识任务上.
- 较大的模型在案例挑战中表现相似,但较小的生物医学模型表现明显不佳.
- 一般用途模型在文本生成,质量保证和编码方面获得了更高的分数,而生物医学LLM则表现出更多的幻觉.
结论:
- 生物医学微调并不本质上改善LLM在看不见的医疗任务上的表现.
- 一般目的的LLM通常优于专门的生物医学模型.
- 检索增强生成为临床LLM集成提供了更有前途的战略.
相关概念视频
Improving Translational Accuracy
2.5K
2.5K
Language and Cognition
308
Language serves as a bridge between ideas and communication, influencing how individuals perceive and interact with the world. Psychologists have long debated whether language shapes thought or vice versa. This discussion gained grip with Edward Sapir and Benjamin Lee Whorf in the 1940s, who proposed that language determines thought, a concept known as linguistic determinism. They suggested that the vocabulary and structure of a language influence how its speakers think and perceive reality.
308


