基于BERT和生成的大型语言模型用于检测自杀念头的比较分析:一项绩效评估研究
Adonias Caetano de Oliveira1,2, Renato Freitas Bessa2, Ariel Soares Teles2,3
1Instituto Federal de Educação, Ciência e Tecnologia do Ceará, Fortaleza, Brasil.
Cadernos de saude publica
|November 28, 2024
概括
与其他大型语言模型和BERT变体相比,微软Bing/GPT-4在巴西葡萄牙语文本中检测自杀想法方面表现优异. 这种人工智能的进步显示出对心理健康支持的希望,但需要临床验证.
科学领域:
- 自然语言处理自然语言处理.
- 人工智能在心理健康中的作用
- 计算语言学 计算语言学
背景情况:
- 人工智能 (AI) 显示出在文本中检测自杀念头的潜力.
- 基于BERT的模型在文本分类任务中表现出色.
- 大型语言模型 (LLM) 可以在没有特殊培训的情况下解决查询.
研究的目的:
- 为了比较三个BERT模型变体和三个LLM (Google Bard,微软Bing/GPT-4,OpenAI ChatGPT-3.5) 在识别巴西葡萄牙文文本中的自杀念头方面的表现.
- 评估AI模型在非临床文本环境中的有效性.
主要方法:
- 使用了2,691个非自杀和1,097个自杀想法句子的数据集,其中100个被选择用于测试.
- 在BERT模型中进行了预处理,超参数优化和交叉验证.
- 通过使用零射击提示工程来评估LLM.
主要成果:
- 微软Bing/GPT-4以98%的准确度实现了最高的性能.
- 精心调整的BERT模型显示出强的结果:BERTimbau-Large (96%),BERTimbau-Base (94%) 和BERT-多语言 (87%).
- 与Bing/GPT-4和微调的BERT模型相比,Google Bard (62%) 和OpenAI ChatGPT-3.5 (81%) 的准确性较低.
结论:
- 这些人工智能模型的高回忆能力表明,有可能减少对有风险的个体的错误分类.
- 虽然这些模型有望支持自杀念头检测,但这些模型缺乏对患者监测的临床验证.
- 建议在使用这些人工智能工具来帮助医疗保健专业人员检测自杀念头时谨慎使用.
相关概念视频
Modeling in Therapy
49
Modeling, a key technique in therapy, uses observational learning to help clients acquire and practice new skills by watching therapists demonstrate desired behaviors. This approach, rooted in Albert Bandura's concept of vicarious learning, plays a significant role in therapeutic interventions for various psychological conditions, including social anxiety, ADHD, and depression.
Participant Modeling
Participant modeling involves therapists demonstrating calm and effective behaviors in...
Participant Modeling
Participant modeling involves therapists demonstrating calm and effective behaviors in...
49
Stereotype Content Model
14.0K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
14.0K


