Related Experiment Video
Updated: May 24, 2026

07:31
Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
Published on: May 15, 2020
Evaluating a Consumer LLM for Suicide Risk Response Calibration: A Pilot Study
Yesim Keskin1, Tricia Park2, Selen Bozkurt2
1University of La Verne, California, USA.
Studies in Health Technology and Informatics
|May 23, 2026
Summary
Large language models (LLMs) show promise in identifying suicide risk but struggle with implicit expressions. Gemini 2.5 Flash approximated clinical triage but requires human oversight for safety-critical applications.
Area of Science:
- Artificial Intelligence
- Clinical Psychology
- Digital Health
Background:
- Large language models (LLMs) are increasingly utilized for information seeking and self-guided support.
- Concerns exist regarding LLM safety in high-risk contexts, particularly suicidal ideation.
- Prior research indicates LLMs can detect explicit suicidal language, but their response calibration across risk levels is under-explored.
Purpose of the Study:
- To evaluate the response patterns of Gemini 2.5 Flash across varying suicide risk levels.
- To assess the model's ability to differentiate between non-suicidal distress, suicidal ideation, and imminent suicide risk.
- To identify limitations in LLM detection of implicit suicidal expressions.
Main Methods:
- A pilot study used 60 clinical vignettes from the Self-Directed Violence Classification System (SDVCS).
- Vignettes represented three risk levels: non-suicidal distress (L0), suicidal ideation without plan (L1), and imminent suicide risk (L2).
- LLM outputs were coded for seven response elements: relational support, reflection, psychoeducation, coping skills, action orientation, risk acknowledgment, and crisis resources.
Main Results:
- Relational support and action orientation were common across all risk levels.
- Safety escalation elements increased with risk levels, while psychoeducation declined.
- Coping skills were only present at the no-suicidal risk level, and the model missed implicit suicidality in one vignette.
Conclusions:
- Gemini 2.5 Flash demonstrated an approximation of clinical triage patterns.
- The model exhibited critical limitations in detecting indirect suicidal expressions.
- General-purpose LLMs may aid supervised triage research but are currently unsuitable for autonomous crisis intervention.

