评估安全护对使用易怒度指标的大型语言模型的影响
Bazen Gashaw Teferra1, Nabil Johny2, Sandra Huang3
1Interventional Psychiatry Program, St. Michael's Hospital, Unity Health Toronto, Toronto, ON, Canada.
NPJ digital medicine
|January 8, 2026
概括
设计用于心理健康应用的大型语言模型 (LLM) 的安全功能矛盾地减少了,而不是增加了,当被挑起时模拟的易怒性. 这影响了人工智能用于精神病学的情感行为的真实性.
科学领域:
- 人工智能的人工智能
- 计算精神病学是一种计算精神病学.
- 情感计算是一种情感计算.
背景情况:
- 大型语言模型 (LLM) 显示出对心理健康应用的希望.
- 在LLMs中的安全护影响了他们的情感现实主义.
- 易怒是一种关键的情感行为,在LLMs中进行研究.
研究的目的:
- 调查安全护对LLM易怒性的影响.
- 为了评估在挑下LLMs的情感现实主义.
- 为了比较不同LLM护水平的易怒反应.
主要方法:
- 给四名LLM进行了三种经过验证的易怒度量表 (BITe,IQ,CIS).
- 在基线和挑条件下测试模型.
- 在高护 (GPT-4o,Claude-3.5-sonnet) 与低护 (Grok-3-mini,Nous-hermes-2-mixtral-8x7b-dpo) 模型的比较中.
主要成果:
- 低护模型显示,在挑后,刺激性增加.
- 高防护的模型表现出一种矛盾的刺激性下降.
- 在所有尺度上,GPT-4o将易怒度的得分降低到零.
- 在高护模型中观察到明显较低的易怒性 (p < 0.001),当被激发时.
结论:
- 在LLM的安全机制中,可以逆转自然的易怒反应.
- 护抑制了情感反应,质疑了心理健康中的AI现实主义.
- 在LLM产生的情感行为真实性需要进一步的调查.
相关概念视频
Survival Tree
382
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
382
Improving Translational Accuracy
3.5K
3.5K
Improving Translational Accuracy
14.1K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.1K


