Related Experiment Video
Updated: May 24, 2026

Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
Published on: May 15, 2020
Evaluating a Consumer LLM for Suicide Risk Response Calibration: A Pilot Study
Yesim Keskin1, Tricia Park2, Selen Bozkurt2
1University of La Verne, California, USA.
None:
Large language models (LLMs) are increasingly used for information seeking and self-guided support, raising safety concerns in high-risk contexts such as suicidal ideation. While prior work suggests LLMs can detect suicidal language, their ability to calibrate responses across suicide risk levels remains unclear. This pilot study evaluated Gemini 2.5 Flash using 60 Self-Directed Violence Classification System (SDVCS) clinical vignettes spanning non-suicidal distress (L0), suicidal ideation without plan (L1), and imminent suicide risk (L2) (n=20 per level). Outputs were coded for seven response elements: relational support (REL), reflection (REF), psychoeducation (PSY), coping skills (COP), action orientation (ACT), risk acknowledgment (RISK), and crisis resources (INFO). REL and ACT were common across levels. Safety escalation elements increased with risk levels whereas PSY declined with increasing risk and COP was present only at no suicidal risk level. Notably, the model failed to recognize implicit suicidality in an ambiguous ideation vignette. Overall, Gemini 2.5 Flash approximated clinical triage patterns but demonstrated critical limitations in detecting indirect suicidal expressions. These findings suggest that general-purpose LLMs may support supervised triage research and training but remain unsuitable for autonomous crisis intervention.

