自杀风险评估通过聊天的眼睛GPT-3.5与聊天GPT-4:维涅特研究
Inbar Levkovich1, Zohar Elyoseph2,3
1Oranim Academic College, Faculty of Graduate Studies, Kiryat Tivon, Israel.
JMIR mental health
|September 20, 2023
概括
在评估自杀风险方面,ChatGPT-4显示出与心理健康专业人员相似的承诺,而ChatGPT-3.5则低估了它. 为了临床应用,需要进一步的研究.
科学领域:
- 人工智能在心理健康中的作用
- 临床决策支持系统 临床决策支持系统
- 自然语言处理应用程序
背景情况:
- 像ChatGPT (OpenAI) 这样的大型语言模型在心理健康方面具有理论潜力.
- 为了预防自杀,ChatGPT的实际应用需要经验验证.
研究的目的:
- 为了评估ChatGPT在评估自杀风险方面的准确性,使用感知到的负担和挫败的归属感.
- 将ChatGPT-4与ChatGPT-3.5.5的自杀风险评估性能进行比较.
主要方法:
- 聊天GPT模型评估了一个有不同程度的患者报告的负担和挫败的归属感.
- 评估与心理健康专业人员的评估进行了比较.
- 在2023年6月至7月,使用ChatGPT-3.5 (3月14日) 和ChatGPT-4 (5月24日) 进行了三个评估程序.
主要成果:
- 聊天GPT-4的自杀企图概率评估与心理健康专业人员一致 (平均Z分数为0.01).
- 与专业人士相比,ChatGPT-3.5显著低估了自杀风险 (平均Z分数为-0.83).
- 与专业人士相比,ChatGPT-4在自杀念头和精神疼痛方面得分更高,在性方面得分更低.
结论:
- 聊天GPT-4向专业人士展示了与专业人员相比较的自杀企图概率评估,并提高了识别自杀想法的准确性.
- 聊天GPT-4对心理疼痛的高估需要进一步调查.
- ChatGPT-3.5对自杀风险的低估令人担忧,强调需要仔细的临床实施和进一步研究.
相关概念视频
Self-Presentation: Self-Monitoring and Self-Handicapping
39.1K
People can go to great lengths to protect their self-image and present themselves in ways that they want others to see them. Sociologist Erving Goffman presented the idea that a person is like an actor on a stage. Calling his theory dramaturgy, Goffman believed that we use “impression management” to present ourselves to others as we hope to be perceived. Each situation is a new scene, and individuals perform different roles depending on who is present (Goffman, 1959). Think about...
39.1K
Generalized Anxiety Disorder
155
Generalized Anxiety Disorder (GAD) is a chronic condition characterized by excessive and uncontrollable worry that persists for at least six months, significantly interfering with daily functioning. Unlike situational anxiety, which arises in response to specific stressors, GAD often occurs without a clear cause. Individuals may experience disproportionate worry about work, health, or relationships. For instance, a person might continuously fear poor health despite normal medical evaluations or...
155
Self-Report Tests of Personality
380
Self-report inventories are objective personality assessments that use multiple-choice items or numbered scales, typically ranging from 1 (strongly disagree) to 5 (strongly agree). They are often called Likert scales after Rensis Likert. These inventories are widely used due to their ease of administration and cost-effectiveness. One of the most prominent examples is the Minnesota Multiphasic Personality Inventory (MMPI), initially developed in the 1940s to assess abnormal personality traits.
380


