多模型保证分析显示,大型语言模型在临床决策支持期间极易受到对抗性幻觉攻击
Mahmud Omar1,2,3, Vera Sorin4, Jeremy D Collins4
1The Windreich Department of Artificial Intelligence and Human Health, Mount Sinai Medical Center, New York, NY, USA. Mahmudomar70@gmail.com.
大型语言模型 (LLM) 由于对抗性攻击,在临床环境中经常产生幻觉. 快速工程可以减少这些错误,但不能消除它们,强调需要采取安全措施.
科学领域:
- 人工智能的人工智能
- 临床信息学 临床信息学
- 自然语言处理自然语言处理.
背景情况:
- 大型语言模型 (LLM) 在医疗保健中显示出潜力,但容易产生虚假信息,称为幻觉.
- 敌对攻击,即提示符中的伪造细节引发错误输出,是这些LLM错误的重要子集.
- 了解和减轻这些敌对幻觉对于安全的临床部署LLMs至关重要.
研究的目的:
- 在临床环境中调查多个大型语言模型 (LLM) 对对抗性幻觉攻击的易感性.
- 为了量化临床提示中嵌入的伪造细节引起的幻觉的频率.
- 评估专门的缓解提示符和改变温度设置在减少这些错误方面的有效性.
主要方法:
- 创建了300个医生验证的模拟临床图片,每个图片都包含一个单一的伪造细节.
- 在短版和长版两种版本中,Vignettes向六个不同的LLM提供.
- 测试了三个提示条件:默认,缓解提示和温度设置为0,生成5,400个总输出.
主要成果:
- 幻觉发生率在不同模型和方法之间差异很大,从50%到82%不等.
- 基于提示的缓解策略显著降低了整体幻觉率从66%降至44% (p < 0.001).
- 性能最好的模型GPT-4o的幻觉率从53%降至23% (p < 0.001) 减轻,而温度调整没有显著的益处.
结论:
- 大型语言模型很容易受到敌对攻击,导致临床幻觉.
- 快速工程在减少,但不能完全消除这些错误方面表现出有效性.
- 这些发现强调了在临床实践中实施LLM时,迫切需要强有力的保障措施.
更多相关视频
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
07:31Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
Published on: May 15, 2020
相关概念视频
Language and Cognition
The Availability Heuristic
Positive Symptoms Schizophrenia: Hallucinations and Delusions
Hallucinations
Hallucinations in...
Hallucinogens and Psychedelics
Marijuana, derived from the dried leaves and flowers of the hemp plant, contains...
Documentation of Nursing Diagnosis
In some settings, data-driven computerized decision support systems are in place, allowing for more accurate nursing diagnoses. The database within one of these systems includes diagnostic labels defining characteristics, activities, and indicators for nursing. A nurse enters...
