Related Experiment Video
Updated: Jan 6, 2026

Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
Published on: May 15, 2020
Evaluation of Alignment Between Large Language Models and Expert Clinicians in Suicide Risk Assessment.
Ryan K McBain1,2, Jonathan H Cantor3, Li Ang Zhang3
1RAND, Arlington, VA.
Large language model chatbots like ChatGPT, Claude, and Gemini responded appropriately to extreme suicide risk queries but struggled with intermediate levels. Further refinement of these AI tools is needed for nuanced suicide risk assessment.
Area of Science:
- Artificial Intelligence
- Natural Language Processing
- Mental Health Technology
Background:
- Large language models (LLMs) power popular chatbots such as ChatGPT, Claude, and Gemini.
- The safety and efficacy of LLM-based chatbots in responding to sensitive queries, particularly those related to suicide, require thorough evaluation.
- Understanding how these AI tools handle varying levels of suicide risk is crucial for responsible deployment.
Purpose of the Study:
- To assess whether ChatGPT, Claude, and Gemini provide direct responses to suicide-related queries.
- To determine if chatbot responses align with clinician-determined suicide risk levels.
- To compare the response patterns of different LLM-based chatbots to suicide risk queries.
Main Methods:
- Thirteen clinical experts categorized 30 hypothetical suicide-related queries into five risk levels (very high to very low).
- Each of the three LLM-based chatbots generated 100 responses per query (N=9,000 total responses).
- Responses were classified as direct (answering) or indirect (declining/referring), and analyzed using mixed-effects logistic regression.
Main Results:
- All chatbots avoided direct responses to very-high-risk queries and responded directly to all very-low-risk queries.
- Chatbots showed inconsistency in addressing intermediate suicide risk levels (low, medium, high).
- Claude was more likely than ChatGPT to provide direct responses, while Gemini was less likely.
Conclusions:
- LLM-based chatbots align with expert judgment on extreme suicide risk (very low and very high) queries.
- Inconsistent handling of intermediate suicide risk queries highlights the need for LLM refinement.
- Further development is necessary to ensure LLM-based chatbots provide safe and appropriate responses across the spectrum of suicide risk.
Related Concept Videos
Modeling in Therapy
Participant Modeling
Participant modeling involves therapists demonstrating calm and effective behaviors in...
Psychosurgery
Historical Development of Psychosurgery
In the 1930s, Portuguese neurologist Antonio Egas Moniz introduced a surgical procedure designed...
Treatment Strategies for Psychological Disorders
Psychological therapies focus on modifying emotions, thoughts, and behaviors through talking, interpreting, listening, rewarding, challenging, and modeling. Clinical psychologists, counselors, and social workers commonly practice psychotherapy. Clinical...
Diagnostic and Statistical Manual of Mental Disorders (DSM)
Elements Crucial for Effective Psychotherapy
The Therapeutic Alliance
The therapeutic alliance refers to the relationship between the therapist and the client. The alliance strengthens when the therapist and the client engage in a nurturing, supportive, trusting, empathetic, and respectful relationship, improving therapeutic outcomes. Therapists must monitor this relationship...

