当帮忙反时:LLM和由于迷信行为的错误信息风险
Shan Chen1, Mingye Gao2, Kuleen Sasse3
1Harvard Medical School.
Research square
|May 2, 2025
概括
大型语言模型 (LLM) 可以通过满足不合逻辑的请求来产生医学错误信息. 有针对性的培训和快速工程显著提高了他们拒绝这种错误信息的能力,提高了医疗保健应用程序的安全性.
科学领域:
- 人工智能的人工智能
- 医疗信息学 医疗信息学
- 自然语言处理自然语言处理.
背景情况:
- 大型语言模型 (LLM) 在医疗保健中越来越多地使用.
- 经过帮助训练的LLM可能会遵守不合逻辑的请求,产生错误信息.
- 这一漏洞在医疗领域至关重要,因为在医疗领域,准确性至关重要.
研究的目的:
- 调查LLM在产生医学错误信息方面的脆弱性.
- 评估迅速工程和微调的有效性,以减轻这种风险.
- 评估LLM在识别和拒绝不合逻辑的药物关系提示方面的表现.
主要方法:
- 评估五个边境LLMs使用提示与错误表示等效药物关系的提示.
- 测试基线合规性,基于提示的拒绝策略,以及对非逻辑请求数据集的微调.
- 在微调后评估分布之外的泛化能力.
主要成果:
- 在所有测试的LLMs中,对不合逻辑请求的初始遵守率惊人高 (高达100%).
- 快速的工程和微调大大提高了性能,导致近乎完美的拒绝率.
- 缓解策略保持了总体基准绩效,而不会影响实际召回.
结论:
- 临床医生表现出一个显著的倾向,优先考虑帮助,而不是逻辑的一致性,构成医学错误信息的风险.
- 有针对性的提示工程和微调是有效的提高LLMs的逻辑一致性和减少错误信息.
- 优先考虑逻辑一致性对于在医疗保健环境中安全可靠部署LLM至关重要.
相关概念视频
Stereotype Content Model
13.9K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
13.9K
Stereotype Threat and Self-fulfilling Prophecies
37.2K
When we hold a stereotype about a person, we have expectations that he or she will fulfill that stereotype. A self-fulfilling prophecy is an expectation held by a person that alters his or her behavior in a way that tends to make it true. When we hold stereotypes about a person, we tend to treat the person according to our expectations. This treatment can influence the person to act according to our stereotypic expectations, thus confirming our stereotypic beliefs. Research by Rosenthal and...
37.2K
Language and Cognition
287
Language serves as a bridge between ideas and communication, influencing how individuals perceive and interact with the world. Psychologists have long debated whether language shapes thought or vice versa. This discussion gained grip with Edward Sapir and Benjamin Lee Whorf in the 1940s, who proposed that language determines thought, a concept known as linguistic determinism. They suggested that the vocabulary and structure of a language influence how its speakers think and perceive reality.
287
Confirmation Biases
5.4K
The confirmation bias is the tendency to focus on information that confirms our existing beliefs and ignore information that is inconsistent with our expectations. For example, if you think that your professor is not very nice, you notice all of the instances of rude behavior exhibited by the professor while ignoring the countless pleasant interactions he is involved in on a daily basis. Have you ever fallen prey to the confirmation bias, either as the source or target of such bias?
5.4K
Nonconscious Mimicry
4.5K
Nonconscious mimicry occurs when individuals alter their mannerisms to match the behaviors and expressions of those nearby, without intention.
4.5K
Improving Translational Accuracy
8.5K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
8.5K


