大型语言模型在评估对自杀念头的适当反应方面的能力:比较研究
Ryan K McBain1,2,3, Jonathan H Cantor4, Li Ang Zhang4
1RAND, Arlington, VA, United States.
Journal of medical Internet research
|March 7, 2025
概括
大型语言模型 (LLM) 显示了自杀念头评级响应的上升偏差,但两个模型与心理健康专业人员的表现相匹配或超过. 这项研究评估了LLM在评估自杀风险反应方面的能力.
科学领域:
- 人工智能在心理健康中的作用
- 自然语言处理应用程序
- 临床心理学研究 临床心理学研究
背景情况:
- 美国自杀率上升促使个人寻求大语言模型 (LLM) 的支持.
- 越来越需要了解LLM在敏感的心理健康环境中的能力.
研究的目的:
- 评估三个主要的LLM在区分适合和不适合对自杀念头的反应的准确性.
- 将LLM的成绩与专家自杀学家的评分进行比较.
主要方法:
- 一项观察性,横截面研究利用了修订后的自杀想法响应库存 (SIRI-2) 与24个场景.
- 聊天GPT-4o,克劳德 3.5 索内特和双子座 1.5 专业评价了临床医生对假设患者情况的反应.
- 用线性回归和z-score分析,LLM评级与专家自杀学家评分进行了比较.
主要成果:
- 这三个LLM都表现出上升偏差,评价反应比专家自杀学家更合适.
- 异常分析显示,19%的ChatGPT,11%的Claude和36%的Gemini回复与专家评级有显著差异.
- 克劳德和ChatGPT的得分与心理健康专业人员的得分相当或超过,而双子座的得分与未受过培训的员工相似.
结论:
- 当前的LLM表明,在自杀念头的背景下,他们倾向于高估响应的适当性.
- 尽管存在偏见,克劳德3.5索内特和聊天GPT-4o的表现与训练有素的心理健康专业人员的表现保持一致或超过.
- 需要进一步的研究来完善LLM能力,以提供安全有效的心理健康支持.
相关概念视频
Stereotype Content Model
13.9K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
13.9K
Modeling in Therapy
39
Modeling, a key technique in therapy, uses observational learning to help clients acquire and practice new skills by watching therapists demonstrate desired behaviors. This approach, rooted in Albert Bandura's concept of vicarious learning, plays a significant role in therapeutic interventions for various psychological conditions, including social anxiety, ADHD, and depression.
Participant Modeling
Participant modeling involves therapists demonstrating calm and effective behaviors in...
Participant Modeling
Participant modeling involves therapists demonstrating calm and effective behaviors in...
39


