Related Experiment Video
Updated: May 23, 2025

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
474
Competency of Large Language Models in Evaluating Appropriate Responses to Suicidal Ideation: Comparative Study
Ryan K McBain1,2,3, Jonathan H Cantor4, Li Ang Zhang4
1RAND, Arlington, VA, United States.
Journal of Medical Internet Research
|March 7, 2025
Summary
Large language models (LLMs) show an upward bias in rating responses for suicidal ideation, but two models match or exceed mental health professional performance. This study assessed LLM competency in evaluating suicide risk responses.
Area of Science:
- Artificial Intelligence in Mental Health
- Natural Language Processing Applications
- Clinical Psychology Research
Background:
- Rising suicide rates in the US prompt individuals to seek support from large language models (LLMs).
- There is a growing need to understand the capabilities of LLMs in sensitive mental health contexts.
Purpose of the Study:
- To evaluate the accuracy of three major LLMs in distinguishing appropriate from inappropriate responses to suicidal ideation.
- To compare LLM performance against expert suicidologist ratings.
Main Methods:
- An observational, cross-sectional study utilized the revised Suicidal Ideation Response Inventory (SIRI-2) with 24 scenarios.
- ChatGPT-4o, Claude 3.5 Sonnet, and Gemini 1.5 Pro rated clinician responses to hypothetical patient situations.
- LLM ratings were compared to expert suicidologist scores using linear regression and z-score analysis.
Main Results:
- All three LLMs exhibited an upward bias, rating responses as more appropriate than expert suicidologists did.
- Outlier analysis revealed 19% of ChatGPT, 11% of Claude, and 36% of Gemini responses differed significantly from expert ratings.
- Claude and ChatGPT achieved scores comparable to or exceeding mental health professionals, while Gemini scored similarly to untrained staff.
Conclusions:
- Current LLMs demonstrate a tendency to overestimate response appropriateness in suicidal ideation contexts.
- Despite bias, Claude 3.5 Sonnet and ChatGPT-4o performance aligns with or surpasses that of trained mental health professionals.
- Further research is needed to refine LLM capabilities for safe and effective mental health support.
Keywords:
ChatGPTSuicidal Ideation Response Inventoryartificial intelligencechatbotdepressiondigital healthlarge language modelmental healthsuicidesuicidologistMore Related Videos
Related Concept Videos
Stereotype Content Model
13.9K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
13.9K
Modeling in Therapy
39
Modeling, a key technique in therapy, uses observational learning to help clients acquire and practice new skills by watching therapists demonstrate desired behaviors. This approach, rooted in Albert Bandura's concept of vicarious learning, plays a significant role in therapeutic interventions for various psychological conditions, including social anxiety, ADHD, and depression.
Participant Modeling
Participant modeling involves therapists demonstrating calm and effective behaviors in...
Participant Modeling
Participant modeling involves therapists demonstrating calm and effective behaviors in...
39

